Papers with Natural Language Processing

300 papers
Proceedings of the Thirteenth Workshop on Graph-Based Methods for Natural Language Processing (TextGraphs-13) (D19-53)

Copied to clipboard

Challenge: TextGraphs is a workshop on graph-based methods for natural language processing . the workshop is being organized in conjunction with the 9th International Joint Conference on Natural Language Processing .
Approach: TextGraphs is the 13th edition of the Workshop on Graph-Based Methods for Natural Language Processing . the workshop promotes synergy between GT and natural language processing .
Outcome: the 2013 edition of TextGraphs is being held in conjunction with the 9th International Joint Conference on Natural Language Processing in Hong Kong.
Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (D18-2)

Copied to clipboard

Challenge: 77 submissions were received for the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) 4 of the 73 valid submissions received were either invalid or withdrawn by the authors.
Approach: The volume contains papers from the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) 4 of the 77 submissions were either invalid or withdrawn by the authors.
Outcome: The system demonstrations session included papers from the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) 4 of the 73 valid submissions were either invalid or withdrawn by the authors.
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (2023.emnlp-main)

Copied to clipboard

Challenge: EMNLP 2023 received 4,909 full paper submissions, the largest number to date . 256 papers were desk rejected for various reasons, leaving us with submissions that were fully reviewed .
Approach: EMNLP 2023 received 4,909 full paper submissions, the largest number to date . 256 papers were desk rejected for various reasons, leaving them fully reviewed .
Outcome: EMNLP 2023 received 4,909 full paper submissions, the largest number to date . 256 papers were desk rejected for various reasons, leaving them fully reviewed .
Findings of the Association for Computational Linguistics: EACL 2024 (2024.findings-eacl)

Copied to clipboard

Challenge: EACL 2024 is the first conference to adopt ARR only . 2024 will be the year where we see whether it actually works .
Approach: cnn's anna mccartney is the general chair of the 18th conference of the European Chapter of the Association for Computational Linguistics . 2024 will be the year where we see whether it actually works .
Outcome: the 18th conference of the European Chapter of the Association for Computational Linguistics will be held in london . organizers decided to move the conference from a month-long conference to a year-long one .
Findings of the Association for Computational Linguistics: EMNLP 2024 (2024.findings-emnlp)

Copied to clipboard

Challenge: EMNLP 2024 will be held in a hybrid format, offering attendees the option to join us in person in Miami, Florida, or to participate remotely from anywhere in the world.
Approach: EMNLP 2024 will be held in a hybrid format, offering attendees the option to join in person or remotely from anywhere in the world.
Outcome: EMNLP 2024 will be held in a hybrid format, offering attendees the option to join in person or remotely from anywhere in the world.
Findings of the Association for Computational Linguistics: EMNLP 2023 (2023.findings-emnlp)

Copied to clipboard

Challenge: EMNLP 2023 received 4,909 full paper submissions, the largest number to date . 256 papers were desk rejected for various reasons, leaving us with submissions that were fully reviewed .
Approach: EMNLP 2023 received 4,909 full paper submissions, the largest number to date . 256 papers were desk rejected for various reasons, leaving them fully reviewed .
Outcome: EMNLP 2023 received 4,909 full paper submissions, the largest number to date . 256 papers were desk rejected for various reasons, leaving them fully reviewed .
Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (2020.emnlp-demos)

Copied to clipboard

Challenge: 91 submissions were received, 10 of which were either invalid or withdrawn .
Approach: 91 submissions were received for the system demonstrations session . 10 were either invalid or withdrawn by the authors .
Outcome: The system demonstrations session was held at the 2020 conference on empirical methods in natural language processing . 91 submissions were accepted, 10 of which were either invalid or withdrawn .
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers) (N18-2)

Copied to clipboard

Challenge: NAACL HLT 2018 is the biggest NAAPL conference to date . this year's conference highlights the vibrancy and vitality of the field .
Approach: a new review form and an opportunity for authors to review the reviewers were introduced at this year's conference . the test-of-time awards are named in memory of Aravind Joshi, who died this year .
Outcome: the biggest NAACL conference to date features a new review form and the Test-of-Time awards . the industrial track features papers that focus on scalable, interpretable, reliable and customer facing methods for industrial applications .
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts (2024.emnlp-tutorials)

Copied to clipboard

Challenge: EMNLP 2024 will feature tutorials on six exciting topics . the process of selecting tutorials was a collaborative effort .
Approach: EMNLP 2024 will feature tutorials on six exciting topics . the process of calling for, submitting, reviewing tutorials was a collaborative effort .
Outcome: the tutorials will cover topics such as natural language explanations, offensive speech, human-centered evaluation, AI for science, agents, and enhancing capabilities of LLMs.
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (2024.emnlp-demo)

Copied to clipboard

Challenge: EMNLP 2024 conference on Empirical Methods in Natural Language Processing received 153 submissions . 52 submissions were selected for inclusion in the program (acceptance rate of 34%)
Approach: EMNLP 2024 conference on Empirical Methods in Natural Language Processing will take place in london on november 12-16, 2024 .
Outcome: The EMNLP 2024 conference is a hybrid event with demonstration papers presented through pre-recorded talks and in presence during the poster sessions.
Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers) (2024.eacl-short)

Copied to clipboard

Challenge: EACL 2024 is the first conference to adopt ARR only . 2024 will be the year where we see whether it actually works .
Approach: cnn's anna mccartney is the general chair of the 18th conference of the European Chapter of the Association for Computational Linguistics . 2024 will be the year where we see whether it actually works .
Outcome: the 18th conference of the European Chapter of the Association for Computational Linguistics will be held in london . organizers decided to move the conference from a month-long conference to a year-long one .
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (2025.emnlp-demos)

Copied to clipboard

Challenge: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing are now available online.
Approach: EMNLP 2025 conference on empirical methods in natural language processing held in Suzhou, china, on November 4-9, 2025. 77 papers accepted for inclusion in proceedings, resulting in 38% acceptance rate.
Outcome: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing are published . the conference accepted 77 papers, with a 38% acceptance rate .
Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (2020.aacl-main)

Copied to clipboard

Challenge: asian-pacific chapter of AACL is hosting its first conference in 2020 . a face-to-face physical meeting would have been eye-opening to participants .
Approach: ai chiang is the General Chair of the Asia-Pacific Chapter of the Association for Computational Linguistics . he is also the General chair of the 10th International Joint Conference on Natural Language Processing .
Outcome: the 1st Asia-Pacific Chapter of the Association for Computational Linguistics will hold its annual conference in 2020 . the conference will be held in conjunction with the 10th International Joint Conference on Natural Language Processing .
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) (N18-1)

Copied to clipboard

Challenge: NAACL HLT 2018 is the biggest NAAPL conference to date . this year's conference highlights the vibrancy and vitality of the field .
Approach: a new review form and an opportunity for authors to review the reviewers were introduced at this year's conference . the test-of-time awards are named in memory of Aravind Joshi, who died this year .
Outcome: the biggest NAACL conference to date features a new review form and the Test-of-Time awards . the industrial track features papers that focus on scalable, interpretable, reliable and customer facing methods for industrial applications .
NLP Lean Programming Framework: Developing NLP Applications More Effectively (N18-5)

Copied to clipboard

Challenge: NLPf is a framework for creating custom natural language processing models and pipelines by utilizing common software development build systems.
Approach: They propose a framework for creating custom NLP models and pipelines by utilizing common software development build systems.
Outcome: This framework allows developers to train and integrate domain-specific NLP pipelines into their applications seamlessly.
Distributed Knowledge Based Clinical Auto-Coding System (P19-2)

Copied to clipboard

Challenge: Codification of free-text clinical narratives has long been recognised to be beneficial for secondary uses such as funding, insurance claim processing and research.
Approach: They propose to use NLP and related machine learning techniques to assign ICD-10-AM and ACHI codes to clinical records using local and international standards.
Outcome: The proposed system utilises NLP and ML techniques to assign ICD-10-AM and ACHI codes to clinical records while adhering to local and international standards.
Mining, Assessing, and Improving Arguments in NLP and the Social Sciences (2023.eacl-tutorials)

Copied to clipboard

Challenge: a tutorial on argument quality assessment will focus on what makes an argument good or bad . argument quality is a field encompassing varying tasks on the automated analysis and synthesis of natural language arguments.
Approach: This tutorial will focus on the assessment of argument quality across disciplines . authors will involve participants in annotation studies on the quality assessment .
Outcome: The tutorial will focus on the assessment of argument quality across disciplines . it will involve participants in two annotation studies on the quality assessment and the improvement of quality .
Towards Generation and Recognition of Humorous Texts in Portuguese (2023.eacl-srw)

Copied to clipboard

Challenge: This PhD thesis focuses on the automatic generation and recognition of verbal punning humor in Portuguese.
Approach: They propose to combine natural language generation and cognitive processing to generate and recognize verbal humor in Portuguese.
Outcome: The proposed methods aim to generate and recognize humor in Portuguese, an underdeveloped language compared to English.
Negation typology and general representation models for cross-lingual zero-shot negation scope resolution in Russian, French, and Spanish. (2021.naacl-srw)

Copied to clipboard

Challenge: Negation resolution remains an acute and continuously researched question in Natural Language Processing.
Approach: They propose to use multilingual pre-trained general representation models to detect negation scope in languages without annotated data.
Outcome: The proposed model achieves token-level F1 score between English, Spanish, French, and Russian.
Discourse Analysis and Its Applications (P19-4)

Copied to clipboard

Challenge: Discourse processing is a suite of NLP tasks to uncover linguistic structures from texts at several levels, which can support many downstream applications.
Approach: They present a set of tasks to uncover linguistic structures from texts at several levels, which can support many downstream applications.
Outcome: The tutorial covers the basic concepts of discourse analysis and linguistic structures in monologue vs. conversation, synchronous v. asynchronous conversation, and key linguistic structure in discourse analysis.
EasyNLP: A Comprehensive and Easy-to-use Toolkit for Natural Language Processing (2022.emnlp-demos)

Copied to clipboard

Challenge: Pre-Trained Models (PTMs) have reshaped the development of natural language processing (NLP) but it is not easy to obtain high-performing PTMs without a large amount of labeled training data and deploy them online with fast inference speed.
Approach: They propose to make it easy to build NLP applications with knowledge-enhanced pre-training and knowledge distillation.
Outcome: EasyNLP supports a comprehensive suite of NLP algorithms and features knowledge-enhanced pre-training, knowledge distillation and few-shot learning functionalities.
Transfer Learning in Natural Language Processing (N19-5)

Copied to clipboard

Challenge: supervised machine learning is based on learning in isolation, a single predictive model for a task using a dataset.
Approach: They present an overview of modern transfer learning methods in natural language processing . they review examples and case studies on how models can be integrated and adapted .
Outcome: The proposed methods improve upon the state-of-the-art on a wide range of NLP tasks.
Can LLMs Learn Macroeconomic Narratives from Social Media? (2025.findings-naacl)

Copied to clipboard

Challenge: Existing evaluation strategies for analyzing economic data with narratives are limited due to the complexity of the interplay of numerous factors and the difficulty in isolating causal relationships.
Approach: They propose to use two Twitter datasets to capture economy-related narratives and use them to construct models using large language models.
Outcome: The proposed models are able to predict macroeconomic fluctuations using the extracted or extracted narratives in two Twitter datasets.
ABLE: Agency-BeLiefs Embedding to Address Stereotypical Bias through Awareness Instead of Obliviousness (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies in Natural Language Processing (NLP) have unveiled a concerning issue: stereotypical biases associated with demographic groups are prevalent.
Approach: They propose an approach that actively encodes stereotypical biases into the embedding space by integrating stereotypes into a model that acquires agency and belief scores rather than directly representing stereotypes.
Outcome: The proposed model can learn agency and belief stereotypes while preserving the language model’s proficiency.
Demo Application for the AutoGOAL Framework (2020.coling-demos)

Copied to clipboard

Challenge: AutoGOAL is a framework for automatically finding the best way to solve a given computational task.
Approach: They present a web demo that showcases the main characteristics of the AutoGOAL framework in Python and a graph-based representation for machine learning pipelines.
Outcome: The proposed framework can be applied to Natural Language Processing and structured classification problems.
Learning with Limited Text Data (2022.acl-tutorials)

Copied to clipboard

Challenge: Natural Language Processing (NLP) relies on labeled data to perform state-of-the-art performance . labeles are often required to label large amounts of textual data . this tutorial will provide an overview of labeleing in NLP .
Approach: This tutorial will provide a systematic overview of methods for learning from limited labeled data.
Outcome: This tutorial will provide a systematic and up-to-date overview of the proposed methods . it will highlight current challenges and future directions .
Mining, Assessing, and Improving Arguments in NLP and the Social Sciences (2024.lrec-tutorials)

Copied to clipboard

Challenge: a tutorial on computational argumentation is updated to address the problem of argument quality . argument quality is a field of interdisciplinary research that connects natural language processing to social sciences .
Approach: They present an updated version of the EACL 2023 tutorial on argument quality . they will focus on the notions of argument quality across disciplines .
Outcome: The updated version of the EACL 2023 tutorial focuses on argument quality assessment . the authors will focus on the interface between Argument Mining and Deliberation Theory .
A Tour of Explicit Multilingual Semantics: Word Sense Disambiguation, Semantic Role Labeling and Semantic Parsing (2022.aacl-tutorials)

Copied to clipboard

Challenge: a recent advent of pretrained language models has sparked a revolution in NLP . but, there are still questions about whether current approaches capture explicit, symbolic meaning . this tutorial will review efforts to tackle three key open problems in lexical and sentence-level semantics .
Approach: This tutorial reviews recent efforts to shed light on meaning in NLP . it will focus on three key open problems in lexical and sentence-level semantics .
Outcome: This tutorial reviews recent efforts to shed light on meaning in NLP . it focuses on three key open problems in lexical and sentence-level semantics .
Graph-based Deep Learning in Natural Language Processing (D19-2)

Copied to clipboard

Challenge: This tutorial aims to introduce graph-based deep learning techniques such as Graph Convolutional Networks (GCNs) for Natural Language Processing (NLP)
Approach: It provides a brief introduction to graph-based deep learning techniques such as Graph Convolutional Networks (GCNs) for Natural Language Processing (NLP).
Outcome: This tutorial provides a brief introduction to graph-based deep learning techniques such as Graph Convolutional Networks (GCNs) for natural language processing (NLP).
CFO: A Framework for Building Production NLP Systems (D19-3)

Copied to clipboard

Challenge: Using a new orchestration framework, we build, test, and deploy interactive NLP and IR systems to production environments.
Approach: They introduce a new orchestration framework for building, experimenting with, and deploying interactive NLP and IR systems to production environments.
Outcome: The proposed framework is well suited to a variety of use cases but is not suitable for academic benchmarking or industry specific use cases.
The iRead4Skills Intelligent Complexity Analyzer (2025.emnlp-demos)

Copied to clipboard

Challenge: 20% of EU adult population exhibits low-literacy and numeracy skills (EA, 2021).
Approach: iRead4Skills Intelligent Complexity Analyzer integrates a range of NLP components to assess input texts along multiple levels of granularity and linguistic dimensions in Portuguese, Spanish, and French.
Outcome: The system assigns four tailored difficulty levels and introduces four diagnostic yardsticks—textual structure, lexicon, syntax, and semantics—offering users actionable feedback on specific dimensions of textual complexity.
Deep Reinforcement Learning for NLP (P18-5)

Copied to clipboard

Challenge: Many natural language processing tasks can be formulated as deep reinforcement learning (DRL) problems.
Approach: This tutorial provides an introduction to the foundations of deep reinforcement learning . it describes recent advances in designing deep reinforcement for NLP .
Outcome: This tutorial provides an introduction to the foundations of deep reinforcement learning and some practical solutions for NLP tasks.
Autodive: An Integrated Onsite Scientific Literature Annotation Tool (2023.acl-demo)

Copied to clipboard

Challenge: Annotating scientific literature directly on PDF documents can greatly improve the labeling efficiency of scientists whose annotation costs are very high.
Approach: They propose an integrated onsite scientific literature annotation tool for natural scientists and Natural Language Processing (NLP) researchers.
Outcome: The proposed tool supports the whole lifecycle of corpus generation including i)project management, ii)resource management, and iv)ontology management, as well as manual annotation, onsite auto annotation, and vi)task statistic.
At a Glance: The Impact of Gaze Aggregation Views on Syntactic Tagging (D19-64)

Copied to clipboard

Challenge: Recent work uses gaze data at the type level or at the token level and mostly from a single eye-tracking corpus.
Approach: They propose to use gaze data to capture central tendency or variability of gaze data and to integrate binary phrase chunking and part-of-speech tagging.
Outcome: The proposed approaches capture the central tendency or variability of gaze data better than proposed local views which retain individual participant information.
An Embarrassingly Simple Approach for Intellectual Property Rights Protection on Recurrent Neural Networks (2022.aacl-main)

Copied to clipboard

Challenge: Existing protection schemes for deep neural network models protect intellectual property rights from being abused, stolen and plagiarized.
Approach: They propose a practical approach for the IPR protection on recurrent neural networks without all the bells and whistles of existing IPR solutions.
Outcome: The proposed approach is robust and effective against ambiguity and removal attacks on different RNN variants.
Pingan Smart Health and SJTU at COIN - Shared Task: utilizing Pre-trained Language Models and Common-sense Knowledge in Machine Reading Tasks (D19-60)

Copied to clipboard

Challenge: Existing approaches to represent knowledge in the low-dimensional space are to leverage large-scale unsupervised text corpus to train fixed or contextual representations.
Approach: They propose to leverage large-scale unsupervised text corpus to train fixed or contextual language representations and to express knowledge into a knowledge graph (KG) they incorporate distributional representations of a KG onto the representations from pre-trained language models, via simply concatenation or multi-head attention.
Outcome: The proposed models outperform the other models on the COIN: COmmonsense INference in Natural Language Processing (COIN) Workshop datasets.
RAGthoven: A Configurable Toolkit for RAG-enabled LLM Experimentation (2025.coling-demos)

Copied to clipboard

Challenge: Large Language Models (LLMs) have significantly altered the landscape of Natural Language Processing (NLP), but their use as a baseline method has not been extensive.
Approach: They propose a tool for automatic evaluation of RAG-based pipelines that provides a simple yet powerful abstraction.
Outcome: The proposed tool provides an automatic evaluation of RAG-based pipelines.
Hard and Soft Evaluation of NLP models with BOOtSTrap SAmpling - BooStSa (2022.acl-demo)

Copied to clipboard

Challenge: Developing better methods for a task is a common feature of the computational linguistics literature.
Approach: They propose to use bootstrap to compute significance levels with the BOOtSTrap SAmpling procedure to evaluate models that predict hard labels and soft labels as well.
Outcome: The proposed method can be used to evaluate models that predict hard labels and soft labels on benchmark data sets.
Establishing Trustworthiness: Rethinking Tasks and Model Evaluation (2023.emnlp-main)

Copied to clipboard

Challenge: Language understanding is a multi-faceted cognitive capability, which the Natural Language Processing community has striven to model computationally for decades.
Approach: They propose to rethink what constitutes tasks and model evaluation in NLP and pursue a more holistic view on language, placing trustworthiness at the center.
Outcome: The proposed models are based on generative models and are being deployed in more real-world scenarios, including previously unforeseen zero-shot setups.
Thesis Proposal: Detecting Agency Attribution (2024.eacl-srw)

Copied to clipboard

Challenge: 'agency' is the freedom and capacity of an entity to act, and the corresponding Natural Language Processing (NLP) task involves automatically detecting attributions of agency to entities in text.
Approach: They propose a schema to annotate a dataset for agency attribution and formulate additional research questions by applying NLP models.
Outcome: The proposed framework draws on semantic frame analysis, role labelling and related techniques.
Monitoring Hate Speech in Indonesia: An NLP-based Classification of Social Media Texts (2024.emnlp-demo)

Copied to clipboard

Challenge: a lack of mechanisms to track the spread and severity of hate speech complicates the formulation of effective solutions.
Approach: They have developed a universally robust hate speech classifier tailored for a narrower subset of texts that target vulnerable groups that have historically been the targets of hate speech in Indonesia.
Outcome: The proposed tool has persuaded the General Election Supervisory Body in Indonesia (BAWASLU) to collaborate with the Alliance of Independent Journalists (AJI) to monitor hate speech in vulnerable areas in the country known for hate speech dissemination or hate-related violence in the upcoming Indonesian regional elections.
Language Technologies for the Creation of Multilingual Terminologies. Lessons Learned from the SSHOC Project (2022.lrec-1)

Copied to clipboard

Challenge: Language Technologies can help in promoting and facilitating multilingualism in the Social Sciences and Humanities domain.
Approach: They propose to use Natural Language Processing and Machine Translation to provide tools to foster multilingual access and discovery to SSH content across different languages.
Outcome: The proposed tools prove to be a valid asset to translation tasks . validation of results by domain experts proficient in the language is an unavoidable phase of the whole workflow.
Zuo Zhuan Ancient Chinese Dataset for Word Sense Disambiguation (2022.naacl-srw)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a core task in natural language processing . ancient Chinese has rarely been used in WSD tasks due to lack of a dataset .
Approach: They annotate ancient Chinese text Zuo Zhuan using a copyright-free dictionary . they apply a method to find the most appropriate sense in a context using k-NN .
Outcome: The proposed dataset will be available on GitHub.
What’s wrong with your model? A Quantitative Analysis of Relation Classification (2024.starsem-1)

Copied to clipboard

Challenge: A major trend in NLP research aims at designing more sophisticated setups to improve the state-of-the-art (SOTA) on a target task.
Approach: They propose an in-depth analysis suite for Relation Classification to be used for prediction tasks.
Outcome: The proposed model improves over the baseline by >3 Micro-F1 . the proposed model is based on a case study and a preliminary error-guided analysis .
Instance-based Inductive Deep Transfer Learning by Cross-Dataset Querying with Locality Sensitive Hashing (D19-61)

Copied to clipboard

Challenge: Existing methods to train supervised learning models rely on labeled data, which is expensive or impossible to acquire.
Approach: They propose an inductive transfer learning method that can augment learning models by infusing similar instances from different learning tasks in Natural Language Processing domain.
Outcome: The proposed method improves the performance of three major news classification datasets by reducing dependency on labeled data by a significant margin.
DILBERT: Customized Pre-Training for Domain Adaptation with Category Shift, with an Application to Aspect Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for pre-training can be sub-optimal in some cases . for example, aspect extraction tasks require domain and category invariant representations .
Approach: They propose a domain-invariant learning scheme for BERT to fine-tune pre-trained language models on a source domain and then apply it to a different target domain.
Outcome: The proposed scheme improves performance over state-of-the-art models while using fraction of the unlabeled data.
Autonomous Machine Learning-Based Peer Reviewer Selection System (2025.coling-demos)

Copied to clipboard

Challenge: Existing systems that match papers with experts are inefficient and often require long turnaround times.
Approach: They propose an autonomous peer reviewer selection system that employs the natural language processing model to match submitted papers with expert reviewers independently of traditional journals and conferences.
Outcome: The proposed system performs faster and smaller than current models while being more scalable.
Automatic Spelling Correction for Resource-Scarce Languages using Deep Learning (P18-3)

Copied to clipboard

Challenge: Indic languages are resource-scarce and do not have such parallel data due to low volume of queries.
Approach: They propose a sequence-to-sequence deep learning model which trains end-to end for Indic languages, Hindi and Telugu.
Outcome: The proposed model is competitive with existing spell checking and correction techniques for Indic languages.
Data Management Plan (DMP) for Language Data under the New General Da-ta Protection Regulation (GDPR) (L18-1)

Copied to clipboard

Challenge: ELRA proposes its own template for the Data Management Plan, which is being updated to take the new law into account.
Approach: They propose a framework for the data management plan to be updated to take the new law into account and propose how it can be integrated into the DMP to increase transparency and spread good practices .
Outcome: The proposed framework will strengthen certain principles related to the processing of personal data, which will also affect many projects in the field of natural language processing.
On Evaluating and Mitigating Gender Biases in Multilingual Settings (2023.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks and resources for evaluating gender biases in multilingual settings are limited.
Approach: They propose to extend DisCo to different Indian languages using human annotations to evaluate gender biases in multilingual models.
Outcome: The proposed benchmarks and mitigation techniques are extended beyond English to evaluate gender biases in multilingual models.
Let Me Know What to Ask: Interrogative-Word-Aware Question Generation (D19-58)

Copied to clipboard

Challenge: Existing models focus on generating questions based on text and the answer to the generated question.
Approach: They propose a pipelined system that predicts the type of interrogative word to be generated . they also propose qg models that can be used to generate questions based on text .
Outcome: The proposed system improves on the task of QG in SQuAD, improving from 46.58 to 47.69 in BLEU-1, 17.55 to 18.53 in blu-4, 21.24 to 22.33 in METEOR, and 44.53 to 46.94 in ROUGE-L.
Massive Choice, Ample Tasks (MaChAmp): A Toolkit for Multi-task Learning in NLP (2021.eacl-demos)

Copied to clipboard

Challenge: Multi-task learning (MTL) has become a standard repertoire in natural language processing (NLP) it enables neural networks to learn tasks in parallel while leveraging the benefits of sharing parameters.
Approach: They propose a toolkit for fine-tuning contextualized embeddings in multi-task settings.
Outcome: The proposed toolkit supports a variety of natural language processing tasks . it enables neural networks to learn tasks in parallel while leveraging the benefits of sharing parameters.
FEAT-writing: An Interactive Training System for Argumentative Writing (2025.coling-demos)

Copied to clipboard

Challenge: Argumentative writing is a critical skill for academic success, but many students struggle to develop these skills.
Approach: They developed an online system that provides students with automated feedback and exercises for argumentative writing.
Outcome: The proposed system improves argumentative writing quality among native English speakers and english-as-a-foreign-language learners.
Mixture of Length and Pruning Experts for Knowledge Graphs Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing graph neural networks (GNNs) adopt rigid, query-agnostic path-exploration strategies limiting their ability to adapt to diverse linguistic contexts and semantic nuances.
Approach: They propose a mixture-of-experts framework that personalizes path exploration . framework uses length experts that adaptively selects and weights candidate paths . it also uses pruning experts that evaluates candidate path from a complementary perspective .
Outcome: The proposed framework shows superior performance on a diverse benchmark . it uses a mixture of experts that weights and selects path lengths according to query complexity .
DistaLs: a Comprehensive Collection of Language Distance Measures (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing work on how to measure distances between languages has focused on intuition and typological distance.
Approach: They propose a toolkit that provides users with easy access to language distance measures.
Outcome: The proposed toolkit provides easy access to a wide variety of language distance measures.
Improving Sentiment Analysis over non-English Tweets using Multilingual Transformers and Automatic Translation for Data-Augmentation (2020.coling-main)

Copied to clipboard

Challenge: Existing models for sentiment analysis over tweets require a substantial amount of text to adapt to a domain where the syntax is different.
Approach: They propose to use a multilingual transformer model to train over tweets in five different languages to adapt the model to non-English languages.
Outcome: The proposed model improves over small corpora of tweets in non-English languages.
MPRF: Interpretable Stance Detection through Multi-Path Reasoning Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Existing stance detection methods treat the task as a classification problem, where models output a stance label without providing interpretable reasoning paths.
Approach: They propose a framework that generates, evaluates, and integrates multiple reasoning paths to improve accuracy, robustness, and transparency in stance detection.
Outcome: The proposed framework outperforms existing models on the SEM16, VAST, and PStance datasets and is highly interpretable and reliable.
On Measures of Biases and Harms in NLP (2022.findings-aacl)

Copied to clipboard

Challenge: Recent studies show that natural language processing (NLP) technologies propagate societal biases about demographic groups associated with attributes such as gender, race, and nationality.
Approach: They propose a framework for harms and questions to help practitioners understand biases . they propose measurable measures to detect and mitigate biased groups .
Outcome: The proposed framework provides a framework for harms and questions for practitioners to answer to guide the development of bias measures.
Fine-mixing: Mitigating Backdoors in Fine-tuned Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for defending NLP models against backdoors have ignored the clean weights of PLMs.
Approach: They exploit pre-trained weights to mitigate backdoors in fine-tuned NLP models . they use a fine-mixing technique and an Embedding Purification technique to do the same .
Outcome: The proposed method outperforms baseline mitigation methods on three single-sentence sentiment classification tasks and two sentence-pair classification tasks.
COGEN: Abductive Commonsense Language Generation (2023.acl-short)

Copied to clipboard

Challenge: Existing training methods for NLP models to perform on two main tasks are needed to introduce these capabilities into the field of reasoning.
Approach: They propose a model that integrates commonsense reasoning with contextual filtering to improve the inference.
Outcome: The proposed model outperforms existing models and sets new state-of-the-art in regards to alphaNLI and alphaNGG tasks.
AdminSet and AdminBERT: a Dataset and a Pre-trained Language Model to Explore the Unstructured Maze of French Administrative Documents (2025.coling-main)

Copied to clipboard

Challenge: Pre-trained language models are used to analyze documents but administrative texts are unstructured and do not perform well.
Approach: They propose a French pre-trained language model for the administrative domain . they compare it with a general domain language model and a large language model .
Outcome: The proposed model improves performance on administrative and general domains.
Learning beyond Datasets: Knowledge Graph Augmented Neural Networks for Natural Language Processing (N18-1)

Copied to clipboard

Challenge: Currently, machine learning is limited in scalability and is limited to specific training data.
Approach: They propose to enhance learning models with world knowledge in the form of Knowledge Graph fact triples for natural language processing tasks.
Outcome: The proposed method is highly scalable to the amount of prior information that has to be processed and can be applied to any generic NLP task.
Twitter-Demographer: A Flow-based Tool to Enrich Twitter Data (2022.emnlp-demos)

Copied to clipboard

Challenge: 199 million people communicate on twitter daily, making it essential to study policy and decision-making.
Approach: They propose a flow-based tool to augment Twitter data with additional information about tweets and users.
Outcome: The proposed tool is designed to enhance Twitter data with additional information about tweets and users.
The D-WISE Tool Suite: Multi-Modal Machine-Learning-Powered Tools Supporting and Enhancing Digital Discourse Analysis (2023.acl-demo)

Copied to clipboard

Challenge: The D-WISE Tool Suite addresses limitations of current DH tools due to the ever-increasing amount of heterogeneous, unstructured, and multi-modal data in which discourses of contemporary societies are encoded.
Approach: They propose to use D-WISE Tool Suite to analyze heterogeneous, unstructured, and multi-modal data in the Digital Humanities (DH)
Outcome: The proposed tool leverages state-of-the-art machine learning technologies from Natural Language Processing and Com-puter Vision to ensure its usability for modernDH research.
Query-Efficient Textual Adversarial Example Generation for Black-Box Attacks (2024.naacl-long)

Copied to clipboard

Challenge: Existing black-box attacks require thousands of queries on the target model, making them expensive in real-world applications.
Approach: They propose a new approach that guides word substitutions using prior knowledge from the training set to improve the attack efficiency.
Outcome: The proposed approach reduces query-free attack and guided search attacks by a factor of 10 500 . it improves transferability and generalization by the ensemble of the ABPens in NLP .
e-CARE: a New Dataset for Exploring Explainable Causal Reasoning (2022.acl-long)

Copied to clipboard

Challenge: Existing causal reasoning models only learn to induce empirical causal patterns that are predictive to the label, while human beings seek for deep and conceptual understanding of the causality to explain the observed causal facts.
Approach: They present a human-annotated CAusal REasoning dataset with conceptual explanations of the causality.
Outcome: The presented dataset shows that human-annotated explanations can be useful for promoting the accuracy and stability of causal reasoning models.
Meta-learning Pathologies from Radiology Reports using Variance Aware Prototypical Networks (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for fewshot learning require a large number of in-domain labeled examples for fine tuning.
Approach: They propose to extend the Prototypical Networks for few-shot text classification by replacing Gaussian class prototypes with a regularization term that encourages the examples to be clustered near the appropriate class centroids.
Outcome: The proposed method outperforms baselines on 13 public and 4 internal datasets and detects potential out-of-distribution (OOD) data points during deployment.
GAINER: Graph Machine Learning with Node-specific Radius for Classification of Short Texts and Documents (2024.eacl-long)

Copied to clipboard

Challenge: Recent advances in Graph Machine Learning (GML) have led to the development of numerous models tailored for processing text for various natural language applications.
Approach: They propose a framework called Graph mAchine learnIng with Node-spEcific Radius that is aimed at graph-based NLP.
Outcome: The proposed framework is non-neural and novel for graph-based NLP.
Generating Equation by Utilizing Operators : GEO model (2020.coling-main)

Copied to clipboard

Challenge: Existing neural models that use hand-crafted features are expensive and lack domain-specific knowledge.
Approach: They propose a GEO model that uses operator-based features to generate equations using natural language sentences.
Outcome: The proposed model outperforms state-of-the-art models on two datasets and 82.1% in ALG514.
Embedding Strategies for Specialized Domains: Application to Clinical Entity Recognition (P19-2)

Copied to clipboard

Challenge: Off-the-shelf word embeddings tend to perform poorly on texts from specialized domains such as clinical reports.
Approach: They combine off-the-shelf contextual embeddings with static word2vec embedders trained on a small in-domain corpus built from task data to reach and sometimes outperform representations learned from a large corpus in the medical domain.
Outcome: The proposed embedding strategies outperform representations learned from a large corpus in the medical domain.
Robustness-Aware Word Embedding Improves Certified Robustness to Adversarial Word Substitutions (2023.findings-acl)

Copied to clipboard

Challenge: Embedding interval bound constraint is important for NLP models to be certified robust, but adversarial examples can be crafted by synonym substitutions.
Approach: They propose a triplet loss to train robustness-aware word embeddings for better certified robustness.
Outcome: The proposed method outperforms state-of-the-art certified defense baselines and generalizes well to unseen substitutions.
Synonym-unaware Fast Adversarial Training against Textual Adversarial Attacks (2025.findings-naacl)

Copied to clipboard

Challenge: Existing adversarial defense methods rely on predetermined linguistic knowledge and assume that attackers’ synonym candidates are known, which is often unrealistic.
Approach: They propose a Fast Adversarial Training method that leverages single-step perturbation generation and effective perturbation initialization to improve model robustness without requiring synonym awareness.
Outcome: Experiments show that the proposed method outperforms existing models under character-level and word-level attacks while still maintaining the correct syntax.
SAPGraph: Structure-aware Extractive Summarization for Scientific Papers with Heterogeneous Graph (2022.aacl-main)

Copied to clipboard

Challenge: Abstractive and extractive methods are used to condense long text into concise summaries while retaining essential information.
Approach: They propose to use paper structure to extract paper summaries from long text . they provide a large-scale dataset of COVID-19-related papers .
Outcome: The proposed framework generates more comprehensive and valuable summaries compared to previous work on COVID-19-related papers.
Bringing the State-of-the-Art to Customers: A Neural Agent Assistant Framework for Customer Service Support (2022.emnlp-industry)

Copied to clipboard

Challenge: Creating agent assistants that can help improve customer service support requires inputs from industry users and their customers as well as knowledge of state-of-the-art natural language processing (NLP) technology.
Approach: They propose to combine expertise from academia and industry to build task/domain-specific Neural Agent Assistants with three high-level components for: (1) Intent Identification, (2) Context Retrieval, and (3) Response Generation.
Outcome: The proposed framework is based on three case studies of industry partners who successfully adapt the framework to their unique challenges.
DeepPavlov 1.0: Your Gateway to Advanced NLP Models Backed by Transformers and Transfer Learning (2024.emnlp-demo)

Copied to clipboard

Challenge: Open-source framework for using NLP models is released for non-experts . complexity of building, fine-tuning and deploying state-of-the-art models remains a barrier .
Approach: They present DeepPavlov 1.0, an open-source framework for using NLP models . the framework is based on PyTorch and supports HuggingFace transformers .
Outcome: The DeepPavlov 1.0 framework is designed for practitioners with limited knowledge of NLP/ML.
Linguistically Grounded Analysis of Language Models using Shapley Head Values (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods for probing language models for morphosyntactic constructions are not well understood . language models gain knowledge of grammatical phenomena during pretraining, but exactly how this knowledge is encoded is not well established.
Approach: They propose a method for probing language models via Shapley Head Values . they use a BLiMP dataset to test their method on linguistic constructions based on a Shaply Head Value method .
Outcome: The proposed method can be used to investigate linguistic knowledge in language models . it shows that attention heads responsible for processing related linguistic phenomena cluster together .
TADA : Task Agnostic Dialect Adapters for English (2023.findings-acl)

Copied to clipboard

Challenge: Existing work on dialectal English NLP is task-specific, using manual annotated dialect data, weak supervision, or data augmentation.
Approach: They propose a method for task-agnostic dialect adaptation by aligning non-SAE dialects with task-specific adapters from SAE.
Outcome: The proposed method improves dialectal robustness on 4 dialectal variants of the GLUE benchmark without task-specific supervision.
Enhancing Clinical BERT Embedding using a Biomedical Knowledge Base (2020.coling-main)

Copied to clipboard

Challenge: Domain knowledge is important for building Natural Language Processing (NLP) systems for low-resource settings, such as in the clinical domain.
Approach: They propose a joint method for adding knowledge base information from the Unified Medical Language System (UMLS) into language model pre-training for some clinical domain corpus.
Outcome: The proposed method outperforms existing models on three clinical domain tasks with no knowledge base information.
TutorialBank: A Manually-Collected Corpus for Prerequisite Chains, Survey Extraction and Resource Recommendation (P18-1)

Copied to clipboard

Challenge: TutorialBank is a publicly available dataset that aims to facilitate NLP education and research . a google search of "Natural Language Processing" returns over 100 million hits with papers, tutorials, 1 http://aan.how blog posts, codebases and other related online resources.
Approach: They have manually collected and categorized over 5,600 resources on NLP . they have created a search engine and command-line tool to search the corpus .
Outcome: The tutorial bank dataset is the largest manually-picked corpus of resources intended for NLP education . it includes lists of research topics, relevant resources for each topic, prerequisite relations among topics .
Considerations for meaningful sign language machine translation based on glosses (2023.acl-short)

Copied to clipboard

Challenge: In machine translation, sign language translation based on glosses is becoming more popular . limitations of glossed approaches are not discussed in a transparent manner, and there is no common standard for evaluation.
Approach: They propose to use a gloss-based approach to evaluate machine translation results . they propose to include realistic datasets, stronger baselines and convincing evaluation .
Outcome: The proposed approach is based on a neural gloss translation model.
Compressing Large-Scale Transformer-Based Models: A Case Study on BERT (2021.tacl-1)

Copied to clipboard

Challenge: Popular pre-trained Transformers have improved performance for various NLP tasks by sizable margins, but are too resource-hungry and computation-intensive to suit low-capacity devices or applications with strict latency requirements.
Approach: They present a literature review of the compression of Transformers, focusing on the popular BERT model, which has attracted considerable research attention.
Outcome: The proposed models improve Sentiment analysis, paraphrase detection, machine reading comprehension, question answering, text summarization, and other tasks by sizable margins.
Sanaphor++: Combining Deep Neural Networks with Semantics for Coreference Resolution (L18-1)

Copied to clipboard

Challenge: Coreference resolution is a challenging task in Natural Language Processing . since a few years, the biggest step forward has been made using deep neural networks .
Approach: They propose to improve coreference resolution by adding semantic features to a top-level deep neural network system . they evaluate a shared task dataset and compare it to the state-of-the-art system based on Stanford deep-coref .
Outcome: The proposed system achieves 1.13% gain over the CoNLL 2012 dataset and the state-of-the-art system.
Modeling the Sacred: Considerations when Using Religious Texts in Natural Language Processing (2024.findings-naacl)

Copied to clipboard

Challenge: This paper concerns the use of religious texts in natural language processing (NLP) religious texts are expressions of culturally important values, and machine learning models reproduce cultural values encoded in training data.
Approach: They argue that NLP's use of religious texts raises considerations beyond model biases . authors argue that religious texts are culturally important and are often used by researchers .
Outcome: The proposed method repurposes translations from their original uses and motivations, and raises considerations beyond model biases.
How to Select One Among All ? An Empirical Study Towards the Robustness of Knowledge Distillation in Natural Language Understanding (2021.findings-emnlp)

Copied to clipboard

Challenge: Knowledge Distillation (KD) is a model compression algorithm that helps transfer knowledge in a large neural network into a smaller one.
Approach: They propose a framework to assess adversarial robustness of multiple KD algorithms.
Outcome: The proposed algorithm achieves state-of-the-art on the GLUE benchmark and out-of domain generalization and adversarial robustness compared to competitive methods.
Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond (2022.tacl-1)

Copied to clipboard

Challenge: causality has not had the same importance in natural language processing, says aaron e. smith . he says research on causality in NLP remains scattered across domains without unified definitions .
Approach: They propose to consolidate research on causality in NLP across academic areas . they explore potential uses of causal inference to improve robustness, fairness, interpretability .
Outcome: The proposed method is a unified overview of causal inference for the NLP community.
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? (2024.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations.
Approach: They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Outcome: The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Potential Idiomatic Expression (PIE)-English: Corpus for Classes of Idioms (2022.lrec-1)

Copied to clipboard

Challenge: Potential Idiomatic Expression (PIE) dataset for NLP in English contains over 20,100 samples with almost 1,200 cases of idioms from 10 classes (or senses).
Approach: They present a large Potential Idiomatic Expression (PIE) dataset for Natural Language Processing (NLP) in English.
Outcome: The proposed dataset contains over 20,100 samples with almost 1,200 cases of idioms (with their meanings) from 10 classes (or senses).
CrowdAgent: Multi-Agent Managed Multi-Source Annotation System (2025.emnlp-demos)

Copied to clipboard

Challenge: Recent approaches to annotate data focus on labeling, but lack holistic process control . a novel system that integrates task assignment, data annotation, and quality/cost management is needed .
Approach: They propose a multi-agent system that integrates task assignment, data annotation, and quality/cost management.
Outcome: The proposed system automates human management by using a collaborative multi-agent system.
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey (2026.findings-eacl)

Copied to clipboard

Challenge: Social media platforms such as X (formerly Twitter), Facebook, and Reddit generate user-generated content.
Approach: They propose a framework to assess privacy risks in social media by evaluating vulnerabilities across six dimensions: data collection, preprocessing, visibility, fairness, computational risk, and regulatory compliance.
Outcome: The proposed framework assesses privacy risks across six dimensions . it achieves F1-scores of 0.58–0.84, but incurs 1% - 23% drop under fine-tuning .
A Neural Few-Shot Text Classification Reality Check (2021.eacl-main)

Copied to clipboard

Challenge: Modern few-shot text classification models struggle when the amount of annotated data is scarce.
Approach: They compare neural few-shot classification models with NLP and computer vision models with transformers to test their performance.
Outcome: The proposed models perform almost equally on ARSC dataset, but not on the intent detection task.
OrchestraLLM: Efficient Orchestration of Language Models for Dialogue State Tracking (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are computationally expensive and often require computational resources.
Approach: They propose a routing framework that seamlessly integrates a SLM and an LLM, or-lm, or a LLM into a single framework.
Outcome: The proposed routing framework reduces the computational costs by over 50% in dialogue state tracking tasks.
Analyzing Homonymy Disambiguation Capabilities of Pretrained Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Processing (NLP) but current pretrained language models lack the granularity to perform disambiguation .
Approach: They propose a large-scale resource that leverages homonymy relations to cluster WordNet senses and train Homonymy Disambiguation systems.
Outcome: The proposed model can distinguish homonyms with up to 95% accuracy even without fine-tuning the underlying PLM.
MiniALBERT: Model Distillation via Parameter-Efficient Recursive Transformers (2023.eacl-main)

Copied to clipboard

Challenge: Pre-trained Language Models (LMs) are an integral part of natural language processing but their usability is constrained by computational and time complexity and their increasing size.
Approach: They propose a technique for converting knowledge of fully parameterised LMs into a compact recursive student.
Outcome: The proposed models match the performance of bloated models with negligible performance losses.
Intent Recognition in Doctor-Patient Interviews (2020.lrec-1)

Copied to clipboard

Challenge: Currently, up to 20 percent of patients are misdiagnosed in medical training programs.
Approach: They propose to annotate doctor-patient interviews with intent inventory and information retrieval methods that are robust with respect to small amounts of training data.
Outcome: The proposed models provide baseline performance scores on the data set for further research.
QA Analysis in Medical and Legal Domains: A Survey of Data Augmentation in Low-Resource Settings (2025.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized natural language processing, but their success remains limited to high-resource domains.
Approach: They analyze the coverage and representativeness of specialized-domain QA datasets against large-scale reference datasets.
Outcome: The proposed methods and evaluations highlight the challenges faced by LLMs in low-resource domains.
FAtNet: Cost-Effective Approach Towards Mitigating the Linguistic Bias in Speaker Verification Systems (2022.findings-naacl)

Copied to clipboard

Challenge: Linguistic bias in Deep Neural Network (DNN) based systems is a critical challenge that needs attention.
Approach: They propose to integrate a lightweight embedding with existing NLP systems to mitigate linguistic bias without adaptation.
Outcome: The proposed framework reduces linguistic bias and enhances usability of baselines for twelve languages.
Sort by Structure: Language Model Ranking as Dependency Probing (2022.naacl-main)

Copied to clipboard

Challenge: Existing algorithms for pre-trained language models lack performance indicators for linguistic tasks such as structured prediction.
Approach: They propose to measure the degree to which labeled trees are recoverable from an LM’s contextualized embeddings by probing to rank LMs for parsing dependencies in a given language.
Outcome: The proposed approach predicts the best LM choice 79% of the time using less compute than training a full parser.
Quantifying Synthesis and Fusion and their Impact on Machine Translation (2022.naacl-main)

Copied to clipboard

Challenge: Literature in Natural Language Processing (NLP) typically labels whole language with strict type of morphology, e.g. fusional or agglutinative.
Approach: They propose to quantify morphological typology at the word and segment level by using two indices: synthesis (e.g. analytic to polysynthetic) and fusion (agglutinative to fusional).
Outcome: The proposed method reduces the rigidity of NLP classification claims by measuring morphological diversity at the word and segment level.
Bringing replication and reproduction together with generalisability in NLP: Three reproduction studies for Target Dependent Sentiment Analysis (C18-1)

Copied to clipboard

Challenge: a lack of reproducibility and generalisability is a major threat to scientific development in Natural Language Processing.
Approach: They propose to use a model zoo to document and release language models and published code . they recommend that future replication experiments should consider a variety of datasets .
Outcome: The proposed methods are compared on six English datasets and are based on the results.
Thesis Proposal: Auditing and Mitigating Demographic Bias in Multi-Stage Retrieval Systems for Criminal Justice Applications (2026.acl-srw)

Copied to clipboard

Challenge: racial descriptors alter embedding similarity scores and retrieval rankings, a new study shows . rife-specific biases can displace relevant records outside top-10 results, the study concludes .
Approach: They propose to detect, measure, and mitigate racial bias in NLP systems deployed in criminal justice contexts . they propose to develop and evaluate debiasing techniques, validate synthetic findings on authentic law enforcement data .
Outcome: The proposed research examines how bias propagates across retrieval pipelines . it shows that racial descriptors alter embedding similarity scores and retrieval rankings .
Increasing Coverage and Precision of Textual Information in Multilingual Knowledge Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to generate knowledge graphs are unable to handle non-English textual information.
Approach: They propose a task of automatic Knowledge Graph Completion to bridge the gap between English and non-English textual information.
Outcome: The proposed method bridges the gap between the quantity and quality of textual information between English and non-English languages.
Discovering and Mitigating Indirect Bias in Attention-Based Model Explanations (2024.findings-naacl)

Copied to clipboard

Challenge: Discrimination is the unfair treatment or prejudice directed towards individuals, groups, or certain ideas or beliefs, intentionally or unintentionally.
Approach: They propose an algorithm to detect and mitigate indirect bias in transformer models by leveraging attention explanations.
Outcome: The proposed algorithm shows that it is more accurate than traditional fairness metrics and that it can be used to mitigate bias in transformer models.
Welcome to the Modern World of Pronouns: Identity-Inclusive Natural Language Processing beyond Gender (2022.coling-1)

Copied to clipboard

Challenge: Current modeling of 3rd person pronouns ignores neopronoun phenomena like naive pronounes, which are not (yet) widely established.
Approach: They propose to validate existing and novel approaches for modeling 3rd person pronouns in language technology and validate them through a survey.
Outcome: The proposed model excludes non-binary users, while ignoring gender-specific phenomena.
ProxyLM: Predicting Language Model Performance on Multilingual Tasks via Proxy Models (2025.findings-naacl)

Copied to clipboard

Challenge: Performance prediction is a method to estimate the performance of Language Models (LMs) on various Natural Language Processing (NLP) tasks.
Approach: They propose a task- and language-agnostic framework to predict the performance of Language Models (LMs) using proxy models.
Outcome: The proposed framework outperforms the state-of-the-art in root-mean-square error (RMSE) and other robustness tests on multilingual NLP tasks.
NLP Scholar: A Dataset for Examining the State of NLP Research (2020.lrec-1)

Copied to clipboard

Challenge: Google Scholar is the largest web search engine for academic literature and provides access to rich metadata associated with the papers.
Approach: They extracted citation information from the ACL Anthology (AA) for about 44 thousand NLP papers and identified authors who published at least three papers there.
Outcome: The ACL Anthology (AA) is the largest repository of articles on Natural Language Processing (NLP).
How Much Do Encoder Models Know About Word Senses? (2025.acl-long)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Processing (NLP) however, how well these models inherently disambiguate word senses remains uncertain.
Approach: They evaluate several encoder-only PLMs across WordNet and ODE sense inventories to evaluate their ability to separate word senses without any task-specific fine-tuning.
Outcome: The proposed model outperforms output layer on WordNet and ODE sense inventories by 15 percentage points.
Gated Multi-Task Network for Text Classification (N18-2)

Copied to clipboard

Challenge: Existing approaches to multitask learning share the features without distinguishing the usefulness of the features, generating undesired interference between tasks.
Approach: They propose to introduce a gate mechanism into multi-task CNN and propose a new gated sharing unit which can filter the feature flows between tasks and greatly reduce the interference.
Outcome: The proposed approach can learn selection rules automatically and gain a great improvement over strong baselines.
Investigating Gender Stereotypes in Large Language Models via Social Determinants of Health (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks evaluate biases related to individual social determinants of health (SDoH) but they overlook interactions between these factors and lack context-specific assessments.
Approach: They investigated the relationship between gender and other SDoH in french patient records to determine whether LLMs rely on embedded stereotypes to make gendered decisions.
Outcome: The proposed models can probe stereotypes and make gendered decisions based on the data.
The Context-Dependent Additive Recurrent Neural Net (N18-1)

Copied to clipboard

Challenge: Contextual sequence mapping is one of the fundamental problems in Natural Language Processing (NLP).
Approach: They propose a new family of Recurrent Neural Networks that address contextual sequence mapping . they propose to use contextual signals to control the flow of information .
Outcome: The proposed architecture outperforms existing methods on dialog problem and language model . the proposed architectures are based on a novel family of recurrent neural networks .
Evaluating the Robustness of Neural Language Models to Input Perturbations (2021.emnlp-main)

Copied to clipboard

Challenge: High-performance neural language models have achieved state-of-the-art results on a wide range of NLP tasks, but results for common benchmark datasets often do not reflect model reliability and robustness when applied to noisy, real-world data.
Approach: They propose to implement character-level and word-level perturbation methods to simulate scenarios in which input texts may be slightly noisy or different from the data distribution on which NLP systems were trained.
Outcome: The proposed methods simulate scenarios in which input texts may be slightly noisy or different from the data distribution on which NLP systems were trained.
EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of emotional intelligence in large language models (LLMs) focus on basic sentiment analysis tasks, such as emotion recognition, which is not enough to evaluate LLMs’ overall emotional intelligence.
Approach: They propose a framework for evaluating the emotional intelligence of large language models (LLMs) that includes four distinct tasks: Key Event Recognition, Mixed Event Recognition and Implicit Emotional Recognition.
Outcome: The proposed framework includes four distinct tasks: Key Event Recognition, Mixed Event Recognition and Implicit Emotional Recognition.
The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing (P18-1)

Copied to clipboard

Challenge: Statistical significance testing is a standard statistical tool designed to ensure that experimental results are not coincidental.
Approach: They propose a protocol for statistical significance test selection in NLP setups . they propose he proposes a survey of the most relevant tests to help guide the protocol .
Outcome: The proposed protocol includes a survey of the most relevant tests.
No Simple Answer to Data Complexity: An Examination of Instance-Level Complexity Metrics for Classification Tasks (2025.naacl-long)

Copied to clipboard

Challenge: Understanding data complexity at the instance level has become increasingly important in Natural Language Processing (NLP) and machine learning (ML).
Approach: They empirically examine the relationship between instance-level complexity scores and metric selection for classification tasks.
Outcome: The results show that storing training loss provides similar complexity rankings to other methods, but not demographic fairness, even in downstream predictions.
uniblock: Scoring and Filtering Corpus with Unicode Block Information (D19-1)

Copied to clipboard

Challenge: Existing methods to remove sentences consisting of illegal characters are tedious and repetitive.
Approach: They propose a statistical method to identify illegal characters in natural language processing . they use a fixed-size feature vector to generate a Gaussian mixture model for each sentence .
Outcome: The proposed method can score sentences and filter corpus on clean corpus and improve performance.
Are Your Keywords Like My Queries? A Corpus-Wide Evaluation of Keyword Extractors with Real Searches (2025.coling-main)

Copied to clipboard

Challenge: Keyword Extraction (KE) is essential in Natural Language Processing (NLP) for identifying key terms that represent the main themes of a text.
Approach: They propose to use real query data from Google Trends to evaluate keywords extracted from a text to capture users' top queries.
Outcome: The proposed method can be used with both supervised and unsupervised KE approaches and shows that KeyBERT is the most effective in capturing users’ top queries.
Can Large Language Models Understand Context? (2024.findings-eacl)

Copied to clipboard

Challenge: Existing evaluation methodologies for Large Language Models (LLMs) have been inadequate to evaluate their ability to understand contextual features.
Approach: They propose a benchmark to assess large language models' ability to understand context by adapting existing datasets to suit their evaluation.
Outcome: The proposed model performs better under the in-context learning pretraining scenario than state-of-the-art models.
Learning Word Meta-Embeddings by Autoencoding (C18-1)

Copied to clipboard

Challenge: Existing word embeddings have shown superior performance in numerous Natural Language Processing (NLP) tasks, however, their performances vary significantly across different tasks.
Approach: They propose to combine distributed word embeddings to produce more accurate and complete meta-embeddings of words.
Outcome: The proposed meta-embeddings outperform the state-of-the-art in multiple tasks.
TED-Q: TED Talks and the Questions they Evoke (2020.lrec-1)

Copied to clipboard

Challenge: Evoked questions represent a hitherto unexplored type of linguistic data, promising to open up important new lines of research.
Approach: They propose a method to annotate TED-talks with the questions they evoke and, where available, the answers to these questions.
Outcome: The proposed method is designed to scale up, relying on crowdsourcing by non-expert annotators, with its utility for Natural Language Processing in mind.
NeuroPrune: A Neuro-inspired Topological Sparse Training Algorithm for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Transformer-based Language Models have become ubiquitous in natural language processing due to impressive performance on various tasks.
Approach: They explore how sparsity affects network topology by exploiting mechanisms seen in biological networks . they show that model-agnostic sparsities are performant across diverse NLP tasks .
Outcome: The proposed model-agnostic sparsity approaches are performant and efficient across NLP tasks.
XLTime: A Cross-Lingual Knowledge Transfer Framework for Temporal Expression Extraction (2022.findings-naacl)

Copied to clipboard

Challenge: Temporal Expression Extraction (TEE) is essential for understanding time in natural language.
Approach: They propose a framework for multilingual Temporal Expression Extraction that leverages pre-trained language models to prompt cross-language knowledge transfer from English to non-English languages.
Outcome: The proposed framework outperforms the existing SOTA methods on French, Spanish, Portuguese, and Basque by large margins.
Using Similarity Measures to Select Pretraining Data for NER (N19-1)

Copied to clipboard

Challenge: Existing studies on how to select appropriate data to pretrain word vectors or LMs are lacking.
Approach: They propose to quantify aspects of similarity between pretraining and target data.
Outcome: The proposed measures are good predictors of the usefulness of pretrained models for Named Entity Recognition over 30 data pairs.
Intrinsic Bias Metrics Do Not Correlate with Application Bias (2021.acl-long)

Copied to clipboard

Challenge: a recent survey of bias in natural language processing found that a coreference system makes more errors in an anti-stereotypical coreferent than in a pro-sterereotype one.
Approach: They compare intrinsic and extrinsic bias metrics across hundreds of trained models . they urge researchers to focus on extrindic measures of bias, not easy to measure .
Outcome: a new intrinsic metric and an annotated test set on gender bias in hate speech are tested . authors urge researchers to focus on extrinsic measures of bias, and to make them more feasible .
Globalizing BERT-based Transformer Architectures for Long Document Summarization (2021.eacl-main)

Copied to clipboard

Challenge: Existing approaches to fine-tune a large language model on downstream tasks show several limitations when the target task requires to reason with long documents.
Approach: They propose a hierarchical approach where the input is divided in multiple blocks independently processed by the scaled dot-attentions and combined between the successive layers.
Outcome: The proposed approach performs well on three extractive summarization corpora of scientific papers and news articles.
At the Crossroad of Cuneiform and NLP: Challenges for Fine-grained Part-of-speech Tagging (2024.lrec-main)

Copied to clipboard

Challenge: cuneiform texts are dominated by multiple languages and language families . the most dominant language written in cuniform is the Semitic Akkadian . existing cnl models are not suitable for digital editions of Akkadi .
Approach: They focus on letters written in the Semitic Akkadian, a cuneiform language dominated by cuniform texts . they propose to use pre-trained embeddings, sentence segmentation and cnl to fine-tune language models .
Outcome: The dominant language written in cuneiform is the Semitic Akkadian . the paper examines the input material and tries to initiate a discussion about best-practices .
Towards AMR-BR: A SemBank for Brazilian Portuguese Language (L18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is a recent and prominent meaning representation with good acceptance and several applications in the Natural Language Processing area.
Approach: They propose to build an AMR annotated corpus for Brazilian Portuguese using an alignment-based approach.
Outcome: The proposed corpus is based on the Little Prince book, which went into the public domain and explored some language-specific annotation issues.
Towards a Welsh Semantic Annotation System (L18-1)

Copied to clipboard

Challenge: Automatic semantic annotation of natural language data is an important task in Natural Language Processing.
Approach: They develop a Welsh semantic annotation tool that can be used to analyze Welsh text . it uses Lancaster's USAS semantic classification scheme to tag words with semantic tags .
Outcome: The proposed tool can cover up to 91.78% of words in Welsh text.
Mitigating Gender Bias in Natural Language Processing: Literature Review (P19-1)

Copied to clipboard

Challenge: NLP models propagate and may even amplify gender bias found in text corpora . methods to mitigate gender bias in NLP are relatively nascent .
Approach: They propose to analyze gender bias based on four forms of representation bias and discuss the advantages and drawbacks of existing gender debiasing methods.
Outcome: The proposed methods are based on four forms of representation bias and have advantages and drawbacks.
Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future (2024.findings-acl)

Copied to clipboard

Challenge: Existing work on social intelligence in NLP does not provide a coherent subfield for researchers to analyze and identify research gaps and future directions.
Approach: They build a social AI taxonomy and a data library of 480 NLP datasets to analyze existing datasets and evaluate language models’ performance in different social intelligence aspects.
Outcome: The proposed infrastructure analyzes existing dataset efforts and evaluates language models’ performance in different social intelligence aspects.
Genre Identification and the Compositional Effect of Genre in Literature (C18-1)

Copied to clipboard

Challenge: Literature is artistic and conveys complex themes over the course of very long narratives.
Approach: They propose a method which can work with large literary corpus of texts . they propose 'gutenberg' dataset to perform Genre Identification .
Outcome: The proposed methods improve results in a literature-based task with 200,000 words of literature . the Gutenberg dataset is used to model literary classifications with a high level of fidelity .
T-VEC: A Telecom-Specific Vectorization Model with Enhanced Semantic Understanding via Deep Triplet Loss Fine-Tuning (2025.emnlp-industry)

Copied to clipboard

Challenge: Generic embedding models struggle to represent telecom-specific semantics . specialized terminology and ambiguous terms often limit their utility in retrieval and downstream tasks.
Approach: They propose a domain-adapted embedding model fine-tuned from a gte-Qwen2-1.5B-instruct backbone.
Outcome: The proposed model outperforms MPNet, BGE, Jina and E5 on a custom benchmark . it is open source and has a triplet loss objective .
PRINCE: Prefix-Masked Decoding for Knowledge Enhanced Sequence-to-Sequence Pre-Training (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on injecting noises into the input sequence, but feasibility of injecting them into the decoding sequence remains an open question.
Approach: They propose a pre-training paradigm that integrates knowledge-enhanced decoding with noises in the prefix to strengthen the representation learning of entities that span over multiple input tokens.
Outcome: The proposed model achieves state-of-the-art results on two knowledge-driven data-to-text generation tasks with up to 2% BLEU gains.
Recent Trends in Linear Text Segmentation: A Survey (2024.findings-emnlp)

Copied to clipboard

Challenge: Linear text segmentation is the task of automatically tagging text documents with topic shifts . the task is based on coherence modeling and/or local cues to identify topic boundaries .
Approach: They provide an overview of current advances in linear text segmentation . they highlight limitations of available resources and of the task itself .
Outcome: The proposed task is based on the most recent literature and under-explored research directions.
Simple Yet Powerful: An Overlooked Architecture for Nested Named Entity Recognition (2022.coling-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is an important task in Natural Language Processing that aims to identify text spans belonging to predefined categories.
Approach: They propose to revisit the Multiple LSTM-CRF (MLC) model, a simple, overlooked, yet powerful approach based on training independent sequence labeling models for each entity type.
Outcome: The proposed model achieves state-of-the-art results in the Chilean Waiting List corpus by including pre-trained language models.
A Zero-shot and Few-shot Study of Instruction-Finetuned Large Language Models Applied to Clinical and Biomedical Tasks (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have enabled advances in the field of natural language processing . however, their application and potential are still underexplored .
Approach: They evaluate four state-of-the-art instruction-tuned Large Language Models on 13 NLP tasks in English.
Outcome: The evaluated models outperform state-of-the-art models on 13 real-world clinical and biomedical NLP tasks in English.
Machine Translationese: Effects of Algorithmic Bias on Linguistic Complexity in Machine Translation (2021.eacl-main)

Copied to clipboard

Challenge: Existing studies have shown that existing models amplify biases observed in training data.
Approach: They propose to use MT and NLP to amplify biases observed in training data to investigate how bias amplification might affect language in a broader sense.
Outcome: The proposed model amplifys biases observed in training data and could lead to an artificially impoverished language, the authors show.
MirasText: An Automatically Generated Text Corpus for Persian (L18-1)

Copied to clipboard

Challenge: Natural language processing is one of the most important fields of artificial intelligence.
Approach: They propose to use MirasText to generate Persian text corpus from Persian websites . MiraSText has over 2.8 million documents and over 1.4 billion tokens .
Outcome: The generated corpus has over 2.8 million documents and over 1.4 billion tokens . MirasText has over 800 billion token tokens and more than 300 thousand articles .
BioRo: The Biomedical Corpus for the Romanian Language (L18-1)

Copied to clipboard

Challenge: Biomedical text mining uses linguistic resources available in English, but for other languages such as Romanian, the access to language resources is not straight-forward.
Approach: They present a biomedical corpus of the Romanian language, which is a valuable linguistic asset for biomedically text mining.
Outcome: The proposed corpus will be made publicly available to the biomedical text mining community . the corpus is a reference corpus for the Romanian language .
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices (2020.acl-main)

Copied to clipboard

Challenge: Empirical studies show that MobileBERT is 4.3x smaller and 5.5x faster than BERT_BASE . BERT is one of the largest models ever in NLP, but suffers from heavy model size and high latency .
Approach: They propose a tool to compress and accelerate the popular BERT model by task-agnostic application.
Outcome: The proposed model is 4.3x smaller and 5.5x faster than BERT_BASE . it achieves competitive results on well-known benchmarks .
Benchmarking GPT-4 on Algorithmic Problems: A Systematic Evaluation of Prompting Strategies (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized the field of natural language processing . however, it has been shown that they lack systematic generalization, which allows to extrapolate the learned statistical regularities outside the training distribution.
Approach: They propose to benchmark a LLM with two parameters to find out its performance . they compare it to a variant of the Transformer-Encoder architecture to find the same problem .
Outcome: The proposed model outperforms the previous model on three algorithmic tasks with two parameters.
Robust Conversational Agents against Imperceptible Toxicity Triggers (2022.naacl-main)

Copied to clipboard

Challenge: Existing work to generate adversarial attacks is costly and not scalable . despite the abundance of research in this area, little attention has been given to adversarials .
Approach: They propose an adversarial attack mechanism that mitigates toxic language generation . they propose a defense mechanism that is scalable and can be generalized .
Outcome: The proposed defense is effective at avoiding toxic language generation even against imperceptible toxicity triggers while preserving conversational flow.
Method Entity Extraction from Biomedical Texts (2022.coling-1)

Copied to clipboard

Challenge: Scientific research papers consist of complex keywords and domain-specific terminologies, and new terminologie erupt.
Approach: They find method terminologies in biomedical text using rule-based and machine learning techniques . authors propose to use a silver standard corpus to extract method entities from biomedically text .
Outcome: The proposed method entities can be extracted from biomedical text with reasonable accuracy . the proposed method entity extraction method is based on a rule-based method and a machine learning technique.
Identification of Primary and Collateral Tracks in Stuttered Speech (2020.lrec-1)

Copied to clipboard

Challenge: Disfluency detection is a challenging task because of its different metrics depending on whether the input features are text or speech.
Approach: They propose a framework for disfluency detection inspired by the clinical and the natural language processing perspective together with the theory of performance from (Clark, 1998) . they present a forced-aligned disfluence dataset and propose new audio features inspired by word-based span features.
Outcome: The proposed framework outperforms baselines for speech-based predictions on a forced-aligned disfluency dataset from semi-directed interviews.
A Comprehensive Survey of Contemporary Arabic Sentiment Analysis: Methods, Challenges, and Future Directions (2025.findings-naacl)

Copied to clipboard

Challenge: Existing literature on Arabic sentiment analysis is limited, compared to high-resourced languages such as English and French.
Approach: They present a systematic review of existing literature on Arabic sentiment analysis focusing on research utilizing deep learning.
Outcome: The proposed methods highlight gaps in the literature on Arabic sentiment analysis and outline promising directions for future research.
Importance Estimation from Multiple Perspectives for Keyphrase Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Existing keyphrase extraction methods focus on the part of phrase that is important . experimental results show that KIEMP outperforms existing keyphrase extracting methods .
Approach: They propose to estimate the importance of keyphrase from multiple perspectives using a chunking module, ranking module and matching module.
Outcome: The proposed method outperforms the state-of-the-art keyphrase extraction methods on six benchmark datasets.
Domain Adaptation for Arabic Cross-Domain and Cross-Dialect Sentiment Analysis from Contextualized Word Embedding (2021.naacl-main)

Copied to clipboard

Challenge: Recent studies have classified dialectal Arabic into more fine-grained levels, including countries and cities.
Approach: They propose to use Arabic domains to transfer knowledge from labeled source domains into unlabeled target domains by transferring the learned knowledge from a labele .
Outcome: The proposed method outperforms other domain adaptation methods and improves performance by 20.8% over the zero-shot transfer learning from BERT.
Document-level Event Factuality Identification via Machine Reading Comprehension Frameworks with Transfer Learning (2022.coling-1)

Copied to clipboard

Challenge: Document-level Event Factuality Identification (DEFI) is a fundamental and crucial task in NLP.
Approach: They propose a framework for document-level event factuality identification (DEFI) they propose to use Span-Extraction and Multiple-Choice to model DEFI as machine reading comprehension tasks .
Outcome: The proposed model outperforms state-of-the-art models on a document-based event factuality task . it uses Span-Extraction (Ext) and Multiple-Choice (Mch) knowledge to extract knowledge from large-scale MRC corpus .
Stress Test Evaluation of Transformer-based Models in Natural Language Understanding Tasks (2020.lrec-1)

Copied to clipboard

Challenge: Existing models are weak and take advantage of failures and errors in datasets to improve performance.
Approach: They evaluate three Transformer-based models in Natural Language Inference and Question Answering tasks to see if they are more robust or have the same flaws as their predecessors.
Outcome: The proposed models outperform recurrent neural network models to stress tests on both NLI and QA tasks.
ILDAE: Instance-Level Difficulty Analysis of Evaluation Data (2022.acl-long)

Copied to clipboard

Challenge: Instance-level difficulty analysis of evaluation data is a new field of research that focuses on leveraging instance difficulty in natural language processing.
Approach: They conduct Instance-Level Difficulty Analysis of Evaluation data in a large-scale setup of 23 datasets and demonstrate its five novel applications.
Outcome: The proposed model improves efficiency and accuracy, improves quality and improves Out-of-Domain performance.
TongGu: Mastering Classical Chinese Understanding with Knowledge-Grounded Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capability in Natural Language Processing (NLP), but struggle with Classical Chinese Understanding (CCU) Existing models, including general-purpose and preliminary LLMs, lack the ability to address CCU in data-demanding and knowledge-intensive tasks.
Approach: They propose to use a classical Chinese corpora-based instruction-tuning dataset to unlock the full CCU potential of LLMs.
Outcome: The proposed model unlocks the full CCU potential of LLMs by preserving its foundational knowledge while maintaining redundancy-aware tuning (RAT) and CCU-RAG.
M5 – A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models and their multimodal counterparts have shown significant performance disparities across different languages and cultural contexts.
Approach: They propose to evaluate LLMs on diverse vision-language tasks within a multilingual and multicultural context using M5 benchmark.
Outcome: The proposed benchmarks highlight task-agnostic performance disparities between languages and cultural contexts.
Improving Word Embedding Factorization for Compression Using Distilled Nonlinear Neural Decomposition (2020.findings-emnlp)

Copied to clipboard

Challenge: Word-embeddings are vital components of natural language processing (NLP) but they consume a lot of memory which poses a challenge for edge deployment.
Approach: They propose an embedding compression method based on matrix decomposition and knowledge distillation that initializes weights of pre-trained word-embeddings and fine-tunes end-to-end.
Outcome: The proposed method has higher BLEU score on translation and lower perplexity on language modeling compared to complex, difficult to tune methods.
Academic-Industrial Perspective on the Development and Deployment of a Moderation System for a Newspaper Website (L18-1)

Copied to clipboard

Challenge: a system that supports the moderation of user comments on a large newspaper website is described in this paper.
Approach: They describe an approach and experiences from the development, deployment and usability testing of a natural language processing and information retrieval system that supports the moderation of user comments on a large newspaper website.
Outcome: The proposed system supports the moderation of user comments on a large newspaper website.
Do Prompt Positions Really Matter? (2024.findings-naacl)

Copied to clipboard

Challenge: Prompt-based learning models have a high level of interest due to their ability to perform zero-shot and fewshot tasks.
Approach: They conduct the most comprehensive analysis to date of prompt position for diverse natural language processing tasks.
Outcome: The proposed model is more robust than previous models and is consistent even in instruction-tuned models.
Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where? (2025.emnlp-main)

Copied to clipboard

Challenge: 20% of all papers in the ACL Anthology address social good issues . authors are more likely to do work addressing social good concerns when publishing in venues outside of ACL.
Approach: They use author- and venue-level perspectives to map the landscape of NLP4SG . they find authors are more likely to do work addressing social good concerns outside of ACL .
Outcome: The study analyzes the literature on NLP4SG and its impact on the ACL community . 20% of all papers in the anthology address social good issues, the study finds .
Proto-lm: A Prototypical Network-Based Framework for Built-in Interpretability in Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for interpreting LLMs are post hoc and focus on low-level features and lack of explainability at higher-level text units.
Approach: They propose a prototypical network-based white-box framework that allows LLMs to learn immediately interpretable embeddings during the fine-tuning stage while maintaining competitive performance.
Outcome: The proposed framework can learn interpretable embeddings during the fine-tuning stage while maintaining competitive performance.
Challenges in Pre-Training Graph Neural Networks for Context-Based Fake News Detection: An Evaluation of Current Strategies and Resource Limitations (2024.lrec-main)

Copied to clipboard

Challenge: Graph Neural Networks (GNNs) are used to train neural networks to detect fake news based on context-based methods.
Approach: They propose to combine the two by applying pre-training of Graph Neural Networks (GNNs) in the domain of context-based fake news detection.
Outcome: The proposed methods show that transfer learning does not lead to significant improvements over training a model from scratch in the domain of context-based fake news detection.
RW-KD: Sample-wise Loss Terms Re-Weighting for Knowledge Distillation (2021.findings-emnlp)

Copied to clipboard

Challenge: Knowledge Distillation (KD) is used to compress the pre-training and task-specific fine-tuning phases of large neural language models.
Approach: They propose a sample-wise loss weighting method that re-weights the two losses for each sample.
Outcome: The proposed method outperforms existing methods on 7 datasets of the GLUE benchmark.
Building Static Embeddings from Contextual Ones: Is It Useful for Building Distributional Thesauri? (2022.lrec-1)

Copied to clipboard

Challenge: contextual language models are dominant in the field of Natural Language Processing, but they are not suitable for all uses.
Approach: They propose a method for building word or type-level embeddings from contextual models . they evaluate a large set of English nouns from the perspective of extracting semantic similarity relations .
Outcome: The proposed method can be used to build word or type embeddings from contextual models . it can be exploited for a wide set of English nouns, showing it can improve distributional thesauri .
Residue-Based Natural Language Adversarial Attack Detection (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches to detect adversarial examples for deep learning based systems focus on image embedding feature spaces . however, existing approaches focus on text features, without considering model embeddable spaces.
Approach: They propose a sentence-embedding “residue” detector to identify adversarial examples from embedded feature spaces.
Outcome: The proposed detector outperforms existing model-focused detectors on many tasks.
CHisIEC: An Information Extraction Corpus for Ancient Chinese History (2024.lrec-main)

Copied to clipboard

Challenge: Historical and cultural heritage preservation is an important branch of digital humanities, where the rich tapestry of the past meets the cutting-edge tools of the digital age.
Approach: They present a dataset to evaluate NER and RE tasks in ancient Chinese history . they use four distinct entity types and twelve relation types to identify them .
Outcome: The "Chinese Historical Information Extraction Corpus" is a dataset from 13 dynasties spanning over 1830 years . the dataset encompasses four distinct entity types and twelve relation types .
With More Contexts Comes Better Performance: Contextualized Sense Embeddings for All-Round Word Sense Disambiguation (2020.emnlp-main)

Copied to clipboard

Challenge: Contextualized word embeddings have been used effectively across several tasks in Natural Language Processing, but it is difficult to link them to structured sources of knowledge.
Approach: They propose a semi-supervised approach to producing sense embeddings for the lexical meanings within a lexicon that is comparable to that of contextualized word vectors.
Outcome: The proposed approach outperforms state-of-the-art models in the English Word Sense Disambiguation task and in the multilingual one while training on sense-annotated data in English only.
RANCC: Rationalizing Neural Networks via Concept Clustering (2020.coling-main)

Copied to clipboard

Challenge: Existing models that construct explanations concurrently with classification predictions are opaque.
Approach: They propose a self-explainable model for Natural Language Processing (NLP) text classification tasks . they extract a rationale from the text and use it to predict a concept of interest .
Outcome: The proposed model can be compressed without complicated compression techniques.
Document-Level Event Factuality Identification via Adversarial Neural Network (N19-1)

Copied to clipboard

Challenge: Document-level event factuality identification is crucial for discourse understanding in NLP . identifying document-level factual of events requires comprehensive understanding of documents .
Approach: They propose to construct a corpus annotated with document- and sentence-level event factuality information on English and Chinese texts.
Outcome: The proposed model outperforms baselines on the constructed corpus.
SupCL-Seq: Supervised Contrastive Learning for Downstream Optimized Sequence Representations (2021.findings-emnlp)

Copied to clipboard

Challenge: SupCL-Seq extends contrastive learning from computer vision to sequence classification tasks.
Approach: They propose a supervised alternative to Masked Language Modeling (MLM) that extends contrastive learning to sequence optimization in NLP by altering the dropout mask probability in standard Transformer architectures.
Outcome: The proposed method leads to large gains on the GLUE benchmark, including 6% absolute improvement on CoLA, 5.4% on MRPC, 4.7% on RTE and 2.6% on STS-B.
Automatic Identification of Research Fields in Scientific Papers (L18-1)

Copied to clipboard

Challenge: TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers.
Approach: TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers.
Outcome: The proposed method will help scientists identify geographical territories areas from scientific papers available in digital versions within and outside the ISTEX library.
GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer (2024.naacl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) models are limited to a set of predefined entity types. Large language models (LLMs) can extract arbitrary entities through natural language instructions.
Approach: They propose a model that can identify any type of entity using a transformer encoder.
Outcome: The proposed model outperforms existing models on NER benchmarks on a set of predefined entities.
Is a Question Decomposition Unit All We Need? (2022.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LMs) have achieved state-of-the-art performance on many NLP benchmarks.
Approach: They propose to decompose a hard question into simpler questions that are easier for models to answer.
Outcome: The proposed approach significantly improves model performance (24% for GPT3 and 29% for RoBERTa-SQuAD along with a symbolic calculator) by decomposing a hard question into simpler questions that are easier for models to answer.
FlauBERT: Unsupervised Language Model Pre-training for French (2020.lrec-1)

Copied to clipboard

Challenge: Language models are a key step to achieve state-of-the-art results in many different Natural Language Processing (NLP) tasks.
Approach: They propose to use a language model that is pre-trained on a large and heterogeneous French corpus to train continuous word representations.
Outcome: The proposed model outperforms existing models on a large and heterogeneous French corpus.
Not Far Away, Not So Close: Sample Efficient Nearest Neighbour Data Augmentation via MiniMax (2021.findings-acl)

Copied to clipboard

Challenge: Existing kNN-based augmentation techniques blindly incorporate all samples, but MiniMax-kNN uses a subset of augmented samples to maximize KL-divergence between teacher and student models.
Approach: They propose a semi-supervised approach to augmented data augmentation using kNN.
Outcome: The proposed method outperforms existing kNN-based augmentation techniques on several classification tasks and requires fewer augmented examples and less computation to achieve superior performance.
Cognitive Information Bottleneck: Extracting Minimal Sufficient Cognitive Language Processing Signals (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to extract only task-relevant information from cognitive processing signals are lacking in the field of NLP.
Approach: They propose a method that extracts only task-relevant information from cognitive processing signals.
Outcome: The proposed method outperforms existing methods in compressing cognitive signals and enhances performance on downstream tasks.
Thinking Outside of the Differential Privacy Box: A Case Study in Text Privatization with Language Model Prompting (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on the integration of Differential Privacy (DP) into NLP techniques.
Approach: They propose a method for text privatization leveraging language models to rewrite texts . they examine the usability of DP in NLP and its benefits over non-DP approaches .
Outcome: The proposed method is a novel method for text privatization leveraging language models to rewrite texts.
A Deep Transfer Learning Method for Cross-Lingual Natural Language Inference (2022.lrec-1)

Copied to clipboard

Challenge: Natural Language Inference (NLI) is a crucial task in AI and natural language processing.
Approach: They propose an effective transfer learning approach for cross-lingual NLI . they perform experiments on English-Hindi language pairs in cross-linguistic setting .
Outcome: The proposed model improves the baseline model by 10% over the state-of-the-art model.
Towards Holistic and Automatic Evaluation of Open-Domain Dialogue Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods of open-domain dialogue evaluation are labor-intensive and inefficient.
Approach: They propose to use open-domain dialogues to evaluate different aspects of dialogues using holistic evaluation metrics.
Outcome: The proposed metrics show strong correlations with human judgments.
Leveraging Abstract Meaning Representation for Knowledge Base Question Answering (2021.findings-acl)

Copied to clipboard

Challenge: Existing approaches face challenges including complex question understanding and lack of large end-to-end training datasets.
Approach: They propose a modular knowledge base question answering system that leverages AMR parses for task-independent question understanding.
Outcome: The proposed system achieves state-of-the-art performance on two prominent KBQA datasets based on DBpedia.
Measuring Robustness for NLP (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to evaluate NLP models are limited to news domains and cannot be generalized to other domains.
Approach: They propose a measure of NLP quality based on robustness . they measure consistency of cross-domain accuracy and introduce coefficient of variation and gamma-Robustness based upon human evaluation .
Outcome: The proposed approach shows higher agreement with human evaluation than accuracy scores on ranking machine translation systems.
A Customized Text Sanitization Mechanism with Differential Privacy (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to sanitize texts subject to differential privacy do not work for non-metric semantic similarity measures.
Approach: They propose a customized text sanitization mechanism based on a metric local differential privacy definition.
Outcome: The proposed mechanism achieves better privacy-utility trade-offs than existing mechanisms on benchmark datasets.
Adapting Word Embeddings to New Languages with Morphological and Phonological Subword Representations (D18-1)

Copied to clipboard

Challenge: Existing approaches to generalization to resource-rich languages are difficult . a recent study shows that word representations can be useful in low resource languages .
Approach: They propose two approaches for improving generalization to low-resource languages by adapting continuous word representations using linguistically motivated subword units.
Outcome: The proposed method improves generalization to low resource languages . it requires neither parallel corpora nor bilingual dictionaries and requires no parallel training .
Developing New Linguistic Resources and Tools for the Galician Language (L18-1)

Copied to clipboard

Challenge: Existing resources and tools for the Galician language are lacking for other less-resourced languages, such as statistical tools for lemmatization and Named Entity Recognition.
Approach: They propose to develop a manually revised corpus for POS tagging and lemmatization, and a new manually annotated corpus to train existing statistical tools for the Galician language.
Outcome: The proposed resources include a new corpus for POS tagging and lemmatization, and a manually annotated corpus to handle Named Entity recognition.
Summarize before Aggregate: A Global-to-local Heterogeneous Graph Inference Network for Conversational Emotion Recognition (2020.coling-main)

Copied to clipboard

Challenge: Existing studies focus on modeling emotion influences with utterance-level features, with little attention paid on phrase-level semantic connection between utterrances.
Approach: They propose a two-stage Summarization and Aggregation Graph Inference Network which integrates inference for topic-related emotional phrases and local dependency reasoning over neighbouring utterances in a global-to-local fashion.
Outcome: The proposed model outperforms the state-of-the-art models on three CER benchmark datasets.
We Need to Talk About train-dev-test Splits (2021.emnlp-main)

Copied to clipboard

Challenge: Standard train-dev-test splits used to benchmark multiple models are now used in NLP . comparing multiple versions of the same model on the test data leads to overfitting and "expiration" of test sets.
Approach: They propose to use a tune-set when developing neural network methods to do model picking.
Outcome: The proposed model picker is more robust against the evaluated hyperparameter ranges than the standard split split.
Modeling Multi-Granularity Hierarchical Features for Relation Extraction (2022.naacl-main)

Copied to clipboard

Challenge: Existing work on relation extraction focuses on constructing explicit structured features using knowledge graph and dependency tree.
Approach: They propose a method to extract multi-granularity features based solely on the original input sentences.
Outcome: The proposed method outperforms state-of-the-art models that even use external knowledge on three public benchmarks: SemEval 2010 Task 8, Tacred, and Tacred Revisited.
Curation of Benchmark Templates for Measuring Gender Bias in Named Entity Recognition Models (2024.lrec-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) models are susceptible to gender bias . benchmark datasets are curated specifically for a given NLP task .
Approach: They propose to filter out benchmark templates with a higher probability of detecting gender bias in NER models.
Outcome: The proposed method is based on masked token prediction and tested in English and german using the corresponding fine-tuned BERT base model.
Challenge Dataset of Cognates and False Friend Pairs from Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Cognates are words that have a common etymological origin and can facilitate the Second Language Acquisition (SLA) however, they also pose a challenge to various NLP applications such as Machine Translation and Cross-lingual Sense Disambiguation.
Approach: They create two cognate datasets for twelve Indian languages and use them to generate cognate sets.
Outcome: The proposed datasets are curated using previously available baseline cognate detection approaches and evaluated with the help of lexicographers.
Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks (2024.lrec-main)

Copied to clipboard

Challenge: Attention pruning techniques have been developed to identify and exploit sparseness . previous work has taken pioneering steps to discover and explain the sparsity in attention patterns .
Approach: They propose a framework that observes attention patterns in a fixed dataset and generates a global sparseness mask.
Outcome: The proposed approach saves 90% of computations and maintains quality of results.
Hi-GEC: Hindi Grammar Error Correction in Low Resource Scenario (2025.coling-main)

Copied to clipboard

Challenge: Automated Grammatical Error Correction (GEC) is a scarcely explored low-resource language . a recent study focused on English, but it focused on Hindi, which presents unique challenges due to its complex syntax and intricate morphology.
Approach: They propose to use a human-edited dataset to generate Hindi GEC data . they also investigate round trip translation using diverse languages for the technique .
Outcome: The proposed method outperforms other methods in Hindi, showing that it is highly efficient.
Modeling Cross-Cultural Pragmatic Inference with Codenames Duet (2023.findings-acl)

Copied to clipboard

Challenge: Existing work on pragmatic reasoning tests using simple word reference games with unidentified speakers and listeners, but speakers' sociocultural background shapes their pragmatic assumptions.
Approach: They propose a dataset which operationalizes sociocultural pragmatic inference in a word reference game.
Outcome: The proposed model improves clue-giving and guessing tasks by accounting for background characteristics and the game context.
GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns (2025.findings-acl)

Copied to clipboard

Challenge: Gender rewriting is an NLP task that uses gendered forms to mitigate gender biases.
Approach: They propose a French gender-neutral rewriting system using collective nouns, which are gender-fixed in French.
Outcome: The proposed system detects gendered forms and replaces them with neutral or opposite forms.
IndiSentiment140: Sentiment Analysis Dataset for Indian Languages with Emphasis on Low-Resource Languages using Machine Translation (2024.naacl-long)

Copied to clipboard

Challenge: Existing solutions to bridge the gap between resource-rich and resource-poor languages are being explored.
Approach: They examine the feasibility of machine translation for creating sentiment analysis datasets in 22 Indian languages.
Outcome: The proposed dataset can be used to tackle low-resource challenges in sentiment analysis for Indian languages.
Generalists vs. Specialists: Evaluating Large Language Models for Urdu (2024.findings-emnlp)

Copied to clipboard

Challenge: Urdu is underrepresented in natural language processing, yet it is underserved.
Approach: They compare general-purpose models with special-purpose ones that have been fine-tuned on specific tasks.
Outcome: The proposed models outperform general-purpose models on seven classification and seven generation tasks.
Large Language Models and Causal Inference in Collaboration: A Comprehensive Survey (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown great potential to enhance Natural Language Processing (NLP) models in areas such as predictive accuracy, fairness, robustness, and explainability.
Approach: They evaluate or improve generative Large Language Models from a causal perspective in areas such as reasoning capacity, fairness and safety issues, explainability, and handling multimodality.
Outcome: The proposed models can be used to perform causal relationship discovery and causal effect estimation tasks.
Can Data Diversity Enhance Learning Generalization? (2022.coling-1)

Copied to clipboard

Challenge: a diversity advanced actor-critical reinforcement learning framework is used to improve NLP generalization and accuracy.
Approach: They introduce Diversity Advanced Actor-Critic reinforcement learning framework to improve NLP generalization and accuracy.
Outcome: The proposed framework outperforms domain adaptation and generalization baselines without using any target domain knowledge.
LongLeader: A Comprehensive Leaderboard for Large Language Models in Long-context Scenarios (2025.naacl-long)

Copied to clipboard

Challenge: LongLeader aims to assess different LLMs' long-context comprehension abilities . long-constext comprehension is a key bottleneck for many use cases .
Approach: They propose a leaderboard to assess different LLMs' long-context comprehension abilities . they offer open-source access to the benchmarks and maintain a dedicated website .
Outcome: The proposed model assesses different LLMs on selected benchmarks and provides open-source access to the benchmarks.
ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty Estimation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods show poor performance under Far Boundary (FB) adversarial examples.
Approach: They propose to use a new technique to detect adversarial examples based on data and model uncertainty to outperform existing methods.
Outcome: The proposed method outperforms existing methods by 3.6 and 6.0 AUC points under each scenario.
The First 100 Days: A Corpus Of Political Agendas on Twitter (L18-1)

Copied to clipboard

Challenge: The first 100 days corpus is a curated corpus of the first 100 of the president and senators . political communication has changed dramatically over recent years .
Approach: They analyze the first 100 days of the president and the senators to see differences in their language usage.
Outcome: The corpus analyzes the first 100 days of the president and the senators to see the differences in their language usage.
Full Parameter Fine-tuning for Large Language Models with Limited Resources (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) require massive GPU resources for training.
Approach: They propose a parameter-efficient optimization that fuses the gradient computation and parameter update in one step to reduce memory usage.
Outcome: The proposed method reduces memory usage to 10.8% compared to the standard approach.
Retrieval and Reasoning on KGs: Integrate Knowledge Graphs into Large Language Models for Complex Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have performed impressively in various NLP tasks, but their inherent hallucination phenomena severely challenge their credibility in complex reasoning.
Approach: They propose to integrate explainable Knowledge Graphs (KGs) with LLMs to alleviate hallucinations . they construct subgraphs to enhance the retrieval capabilities of KGs via CoT reasoning.
Outcome: Extensive experiments on two KGQA datasets show that the proposed model achieves convincing performance compared to strong baselines.
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.
JASS: Japanese-specific Sequence to Sequence Pre-training for Neural Machine Translation (2020.lrec-1)

Copied to clipboard

Challenge: Neural machine translation (NMT) requires large parallel corpora for training robust and high quality models.
Approach: They propose a Japanese-specific sequence to sequence pre-training alternative to MASS for NMT . they use Japanese as the source or target language to train their models .
Outcome: The proposed approach can give competitive results over MASS and BRSS, and significantly surpass the individual methods.
CLISTER : A Corpus for Semantic Textual Similarity in French Clinical Narratives (2022.lrec-1)

Copied to clipboard

Challenge: Modern Natural Language Processing relies on the availability of annotated corpora for training and evaluation.
Approach: They propose to annotate sentences in French using a definition of similarity guided by clinical facts and use it to evaluate the corpus.
Outcome: The proposed model can capture similarity with state-of-the-art performance on the DEFT STS shared task evaluation data set.
Continual Few-Shot Learning for Text Classification (2021.emnlp-main)

Copied to clipboard

Challenge: a large number of end-to-end systems are needed for many tasks in natural language processing.
Approach: They propose a continual few-shot learning task where a system is asked to correct mistakes with a few training examples.
Outcome: The proposed task compares two NLI and one sentiment analysis datasets with baselines from diverse paradigms.
CNER: Concept and Named Entity Recognition (2024.naacl-long)

Copied to clipboard

Challenge: Concept and Named Entity Recognition (CNER) is a new unified task that handles concepts and entities mentioned in unstructured texts seamlessly.
Approach: They propose a new unified task that handles concepts and entities mentioned in unstructured texts seamlessly.
Outcome: The proposed task gains +5.4 and +8 macro F1 points when performed as a unified task compared to specialized named entity and concept recognition systems.
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing web crawling pipelines are used to collect large corpora raw data, but the main way to collect such data is through manual data extraction.
Approach: They propose to use a web crawler to extract and classify data from a multilingual web corpus and an automated annotation pipeline to improve it.
Outcome: The proposed version of OSCAR could be used to pre-train large generative language models and other applications in Natural Language Processing and Digital Humanities.
Low-Rank Updates of pre-trained Weights for Multi-Task Learning (2023.findings-acl)

Copied to clipboard

Challenge: Multi-task learning is a popular approach for learning with pre-trained models due to the complexity of the tasks and the challenges associated with fine-tuning large pre-train models.
Approach: They propose a new approach for Multi-task learning which is based on stacking the weights of Neural Networks as a tensor.
Outcome: The proposed approach achieves equivalent performance to the state-of-the-art on the general language understanding evaluation benchmark by training only 0.3 of the parameters per task while not modifying the baseline weights.
An Empirical Study on Explanations in Out-of-Domain Settings (2022.acl-long)

Copied to clipboard

Challenge: Recent work in Natural Language Processing has focused on extracting faithful explanations . yet, little is known about how post-hoc explanations perform in out-of-domain settings .
Approach: They propose to use a random baseline to evaluate out-of-domain post-hoc explanation faithfulness . they suggest select-then-predict models demonstrate comparable predictive performance in out- of-domain settings to full-text trained models.
Outcome: The proposed models perform better in out-of-domain settings than full-text models.
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols .
Approach: They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data .
Outcome: The proposed benchmark assesses pre-trained language models on 20 diversified tasks.
KINNEWS and KIRNEWS: Benchmarking Cross-Lingual Text Classification for Kinyarwanda and Kirundi (2020.coling-main)

Copied to clipboard

Challenge: low-resource African languages are traditionally left behind because of the lack of well-annotated data and effective preprocessing.
Approach: They propose two news datasets for multi-class classification of news articles in two low-resource African languages.
Outcome: The proposed datasets show that training embeddings on the higher-resourced Kinyarwanda yields successful cross-lingual transfer to Kirundi.
BadWindtunnel: Defending Backdoor in High-noise Simulated Training with Confidence Variance (2025.findings-acl)

Copied to clipboard

Challenge: Current backdoor attack defenders in NLP typically involve data reduction or model pruning, risking losing crucial information.
Approach: They propose a backdoor defender that allows precise control over training conditions to model backdoor learning behavior without affecting the final model.
Outcome: The proposed model reduces the backdoor learning behavior without affecting the final model.
Dynamic Stashing Quantization for Efficient Transformer Training (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive performance on a range of Natural Language Processing (NLP) tasks.
Approach: They propose a dynamic quantization strategy that reduces the amount of memory operations and reduces arithmetic cost by 20.95 on two translation tasks and three classification tasks.
Outcome: The proposed model reduces the amount of arithmetic operations by 20.95 and the number of DRAM operations by 2.55 on two translation tasks and three classification tasks.
How do humans perceive adversarial text? A reality check on the validity and naturalness of word-based adversarial attacks (2023.acl-long)

Copied to clipboard

Challenge: Existing text adversarial attacks are impractical in real-world scenarios where humans are involved.
Approach: They have surveyed 378 human participants about the perceptibility of text adversarial examples produced by state-of-the-art methods.
Outcome: The proposed methods ignore the property of imperceptibility or study it under limited conditions.
HECTOR: A Hybrid TExt SimplifiCation TOol for Raw Texts in French (2022.lrec-1)

Copied to clipboard

Challenge: Existing systems for automatic text simplification (ATS) focus on lexical and syntactic transformations, but there is no end-to-end system for French.
Approach: They propose to use word embeddings for lexical simplification and rule-based strategies for syntax and discourse adaptations to improve the complexity of texts.
Outcome: The proposed system performs at lexical, syntactic and discourse levels according to automatic and humanevaluations.
Building an English-Chinese Parallel Corpus Annotated with Sub-sentential Translation Techniques (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that human translators often resort to different non-literal translation techniques besides literal translation . however, they receive less attention in developing natural language processing (NLP) applications.
Approach: They propose to have a better semantic control of extracting paraphrases from bilingual parallel corpora.
Outcome: The proposed method can automatically recognize different non-literal translation techniques . the results confirm the hypothesis of the proposed method .
What do Large Language Models Learn beyond Language? (2022.findings-emnlp)

Copied to clipboard

Challenge: Pretraining on text confers models with useful ‘inductive biases’ for non-linguistic reasoning.
Approach: They investigate whether pre-training on text confers these models with helpful ‘inductive biases’ for non-linguistic reasoning.
Outcome: The proposed models outperform non-pretrained models on 19 non-linguistic tasks and show that they retain inductive biases even when training on multi-lingual text and computer code.
Limitations of Language Models in Arithmetic and Symbolic Induction (2023.acl-long)

Copied to clipboard

Challenge: Recent work has shown that large pretrained Language Models (LMs) can perform remarkably well on a range of NLP tasks but they have limitations on basic symbolic manipulation tasks such as copy, reverse, and addition.
Approach: They propose to use explicit positional markers, fine-grained computation steps, and LMs with callable programs to teach large pretrained Language Models.
Outcome: The proposed model can perform 100% accuracy in OOD and repeating symbols.
Learning Gender-Neutral Word Embeddings (D18-1)

Copied to clipboard

Challenge: Word embeddings trained on human-generated corpora inherit strong gender stereotypes . prior studies show such embeddables exhibit social biases, such as gender stereotype .
Approach: They propose a method to preserve gender information in certain dimensions of word vectors . they propose GN-GloVe, which is a gender-neutral variant of the word embedding model .
Outcome: The proposed method preserves gender information in certain dimensions of word vectors while compelling other dimensions to be free of gender influence.
This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have grammatical knowledge but fail to interpret negation . a recent study shows that LLMs struggle with negative sentences .
Approach: They propose to use a dataset to grasp LLMs' generalization and inference capability . they also fine-tuned models to assess whether the understanding of negation can be trained .
Outcome: The proposed model is able to generalize and infer negation in 400,000 sentences . but it is suboptimal when it comes to negation, a key step in natural language processing .
A Prism Module for Semantic Disentanglement in Name Entity Recognition (P19-1)

Copied to clipboard

Challenge: Xu et al., 2015) proposed a noise reduction mechanism to disentangle semantics of words . hard and soft attention mechanisms are used to reduce noise in NLP tasks .
Approach: They propose a prism module to disentangle semantic aspects of words and reduce noise . they propose combining prism modules with downstream models to improve model performance .
Outcome: The proposed method significantly improves the performance of baselines on named entity recognition (NER) tasks.
Regulation and NLP (RegNLP): Taming Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: polarization in AI safety and ethics debates are swaying political agendas on AI regulation and governance . regulation studies are rich source of knowledge on how to systematically deal with risk and uncertainty .
Approach: They argue that NLP research can benefit from proximity to regulatory studies . they argue that regulation studies should focus on linking scientific knowledge to regulatory processes .
Outcome: The proposed research space should focus on linking scientific knowledge to regulatory processes based on systematic methodologies.
MisinfoBench: A Multi-Dimensional Benchmark for Evaluating LLMs’ Resilience to Misinformation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks assess factual accuracy in isolated queries but fail to evaluate LLMs’ resilience to misinformation in interactive settings.
Approach: MisinfoBench is a benchmark designed to assess LLMs’ ability to discern, resist, and reject misinformation.
Outcome: MisinfoBench assesses large language models’ ability to discern, resist, and reject misinformation in interactive settings.
Dedicated Language Resources for Interdisciplinary Research on Multiword Expressions: Best Thing since Sliced Bread (2020.lrec-1)

Copied to clipboard

Challenge: Multiword expressions are challenging for disciplines like NLP, psycholinguistics and second language acquisition due to their more or less fixed character.
Approach: They propose to develop tools and language resources that are crucial for multifaceted research.
Outcome: The proposed tools and language resources are crucial for this kind of multifaceted research.
HeterGraphLongSum: Heterogeneous Graph Neural Network with Passage Aggregation for Extractive Long Document Summarization (2022.coling-1)

Copied to clipboard

Challenge: Existing models for extractive document summarization are based on sequence-to-sequence (Seq2Sequency) but long-form document summaries using graph-based methods are still an open research issue.
Approach: They propose a heterogeneous graph neural network model to improve the performance of extractive document summarization using graph-based methods.
Outcome: The proposed model can achieve state-of-the-art performance without pre-trained language models.
Human-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech Countering (2022.emnlp-main)

Copied to clipboard

Challenge: a new approach to combat online hate speech is being proposed for NLG . existing methods to train NLG are limited to 2-turn interactions, while in real life, interactions can consist of multiple turns.
Approach: They propose to combine human annotators with machine generated dialogues to create a dataset . DIALOCONAN is the first dataset comprising over 3000 fictitious multi-turn dialogues .
Outcome: The proposed approach combines human experts over machine generated dialogues . it is the first dataset comprising over 3000 fictitious multi-turn dialogues between a hater and an NGO operator .
Do Large Language Models Know What They Don’t Know? (2023.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have vast knowledge that allows them to excel in various NLP tasks.
Approach: They propose an automated method to detect uncertainty in the responses of large language models and a dataset to measure their self-knowledge.
Outcome: The proposed method detects uncertainty in the responses of large language models and provides a novel measure of their self-knowledge.
Sensitive Data Detection and Classification in Spanish Clinical Text: Experiments with BERT (2020.lrec-1)

Copied to clipboard

Challenge: Massive digital data processing can endanger personal data privacy . anonymisation involves removing or replacing sensitive information from data .
Approach: They propose to use a BERT-based sequence labelling model to conduct an experiment on clinical datasets in Spanish.
Outcome: The proposed model outperforms existing models on clinical datasets in Spanish and shows that it is highly competitive with other models.
Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading Comprehension (2024.emnlp-main)

Copied to clipboard

Challenge: Extractive Machine Reading Comprehension (MRC) is a challenging field in the field of Natural Language Processing.
Approach: They propose a Question-Attended Span Extraction module to address the limitations of generative approaches for extractive machine reading comprehension (MRC) . module significantly enhances performance of pre-trained generative language models, enabling them to surpass the extractive capabilities of advanced Large Language Models (LLMs)
Outcome: The QASE module surpasses state-of-the-art models in few-shot settings.
DisGeM: Distractor Generation for Multiple Choice Questions with Span Masking (2024.findings-emnlp)

Copied to clipboard

Challenge: Multiple-choice cloze tests are a prevalent form of assessment that evaluates students' comprehension and inference abilities.
Approach: They propose a framework for distractor generation using readily available pre-trained language models . human evaluations confirm that their approach produces more effective distractors .
Outcome: The proposed framework outperforms existing methods without training or fine-tuning human evaluations confirm it.
Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus (2020.coling-main)

Copied to clipboard

Challenge: Large text corpora are increasingly important for a wide variety of NLP tasks.
Approach: They propose to train automatic language identification models on up to 1,629 languages . they find that human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages.
Outcome: The proposed models achieve over 90% average F1 on 1,629 languages . human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages - suggesting a need for more robust evaluation.
RoBERT – A Romanian BERT Model (2020.coling-main)

Copied to clipboard

Challenge: Existing pre-trained language models learn contextualized representations by using unlabeled text data and obtain state of the art results on a multitude of NLP tasks.
Approach: They propose a pre-trained BERT model for Romanian language processing and compare it with multi-lingual models on seven Romanian specific NLP tasks.
Outcome: The proposed model outperforms multi-lingual models on seven Romanian specific NLP tasks on sentiment analysis, dialect and cross-dialect topic identification, and diacritics restoration.
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics (2025.naacl-long)

Copied to clipboard

Challenge: Recent efforts to improve the quality of machine-generated natural language content have been limited due to the large token usage required by complex evaluation prompts.
Approach: They propose a prompt optimization approach that uses a smaller, fine-tuned language model to compress input data for evaluation prompt, thus reducing token usage and computational cost when using larger LLMs for downstream evaluation.
Outcome: The proposed approach reduces token usage and costs by 2.37 compared with larger LLMs for downstream evaluation.
Word Embedding Evaluation in Downstream Tasks and Semantic Analogies (2020.lrec-1)

Copied to clipboard

Challenge: Language Models (LMs) are an oft studied area of natural language processing . Word Embeddings (WE) are vector space representations of a vocabulary .
Approach: They evaluate Word Embeddings (WE) models for the Portuguese langauage . results show that a diverse corpus can often outperform a larger, less textually diverse corp.
Outcome: The proposed models outperform a larger, less textually diverse corpus in two tasks . the evaluation shows that a diverse and comprehensive corpus outperformed a smaller, less diverse corp.
Automatic Mathematic In-Context Example Generation for LLM Using Multi-Modal Consistency (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for in-context learning require annotated datasets, resulting in higher computational costs and lower quality examples.
Approach: They propose a framework that automatically generates high-quality in-context examples to enhance LLMs’ mathematical reasoning.
Outcome: Evaluated on four math problem datasets, the proposed framework outperforms baseline methods with LLM accuracy ranging from 87.0% to 99.3%.
Adapt or Get Left Behind: Domain Adaptation through BERT Language Model Finetuning for Aspect-Target Sentiment Classification (2020.lrec-1)

Copied to clipboard

Challenge: Aspect-Target Sentiment Classification (ATSC) is a subtask of Aspect Based Sentimence Analysis (ABSA) . recent deep transfer-learning methods have been applied successfully to a myriad of NLP tasks.
Approach: They propose to use a self-supervised domain-specific BERT language model to exploit ATSC . they also perform cross-domain evaluation to explore the real-world robustness of their models .
Outcome: The proposed model outperforms baseline models on the SemEval 2014 task 4 restaurants dataset.
Bridging Robustness and Generalization Against Word Substitution Attacks in NLP via the Growth Bound Matrix Approach (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that adversarial examples can alter models' predicted sentiment due to their sensitivity to specific word choices.
Approach: They propose a regularization technique to improve NLP model robustness by reducing the impact of input perturbations on model outputs.
Outcome: The proposed method outperforms state-of-the-art methods in adversarial defense.
Recontextualizing Revitalization: A Mixed Media Approach to Reviving the Nüshu Language (2025.emnlp-main)

Copied to clipboard

Challenge: Nüshu is an endangered language from Jiangyong County, Hunan, China, and the world’s only known writing system created and used exclusively by women.
Approach: They propose to use NüshuStrokes to record all 397 Unicode Nü Shu characters in sequential handwriting by an expert calligrapher.
Outcome: Evaluating five state-of-the-art Chinese Optical Character Recognition systems on NüshuVision lowers CER to 0.67, a modest but meaningful improvement over previous datasets.
Back to the Basics: A Quantitative Analysis of Statistical and Graph-Based Term Weighting Schemes for Keyword Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Term weighting schemes are widely used in Natural Language Processing and Information Retrieval.
Approach: They perform an exhaustive and large-scale empirical comparison of term weighting methods in the context of keyword extraction using tf-idf.
Outcome: The proposed methods have advantages over tf-idf, and qualitative differences between them.
Vector-Vector-Matrix Architecture: A Novel Hardware-Aware Framework for Low-Latency Inference in NLP Applications (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to improve accuracy of neural networks are slow due to computational complexity.
Approach: They propose a vector-vector-matrix architecture which greatly reduces latency at inference time for NLP applications by a factor of four.
Outcome: The proposed framework reduces the latency of sequence-to-sequence and Transformer models used for NMT by a factor of four.
CamemBERT: a Tasty French Language Model (2020.acl-main)

Copied to clipboard

Challenge: Pretrained language models are now ubiquitous in Natural Language Processing, but their use in other languages is limited.
Approach: They propose to train monolingual Transformer-based model for other languages using web crawled data instead of Wikipedia data and a relatively small web crawl dataset leads to better results.
Outcome: The proposed model performs as well as those obtained using larger datasets.
Building a Sentiment Corpus of Tweets in Brazilian Portuguese (L18-1)

Copied to clipboard

Challenge: Sentiment analysis is a popular area of Natural Language Processing due to its subjective and semantic characteristics.
Approach: They propose to annotate Brazilian Portuguese sentences manually using a sentiment corpus . they run experiments on polarity classification using six machine learning classifiers .
Outcome: The proposed method is based on a Brazilian Portuguese sentiment corpus and achieved 80.38% on F-Measure and 64.87% when including the neutral class.
ESCOXLM-R: Multilingual Taxonomy-driven Pre-training for the Job Market Domain (2023.acl-long)

Copied to clipboard

Challenge: Increasing number of NLP benchmarks highlight need for multilingual models for job-related tasks.
Approach: They introduce a language model called ESCOXLM-R that uses domain-adaptive pre-training on the European Skills, Competences, Qualifications and Occupations taxonomy.
Outcome: The proposed model outperforms XLM-R-large on short spans and entity-level and surface-level span-F1 tasks on entity- and surface level.
Classist Tools: Social Class Correlates with Performance in NLP (2024.acl-long)

Copied to clipboard

Challenge: despite growing concerns surrounding fairness and bias in NLP, there is a dearth of studies delving into the effects it may have on NLP systems.
Approach: They argue that NLP systems’ performance is affected by speakers’ SES, potentially disadvantaging less-privileged socioeconomic groups.
Outcome: The proposed model shows that NLP systems perform better on tasks with social class, ethnicity and geographical variation than those without social class.
What’s the Meaning of Superhuman Performance in Today’s NLU? (2023.acl-long)

Copied to clipboard

Challenge: Recent research has focused on developing larger pretrained language models and introducing benchmarks such as SuperGLUE and SQuAD to measure their abilities.
Approach: They propose to use benchmarks such as SuperGLUE and SQUAD to evaluate PLMs' abilities in language understanding, reasoning, and reading comprehension to assess their performance.
Outcome: The proposed benchmarks have serious limitations affecting comparison between humans and PLMs and provide recommendations for fairer and more transparent benchmarks.
GreekBART: The First Pretrained Greek Sequence-to-Sequence Model (2024.lrec-main)

Copied to clipboard

Challenge: Transfer learning has revolutionized the fields of Computer Vision and Natural Language Processing.
Approach: They introduce a new language model, GreekBART, that is based on a BART-base architecture.
Outcome: The proposed model outperforms BERT, GPT and other transformer-based models on discriminative tasks.
Gender Gap in Natural Language Processing Research: Disparities in Authorship and Citations (2020.acl-main)

Copied to clipboard

Challenge: Disparities in authorship and citations across gender can have adverse consequences . Historically, gender has been considered binary (male and female), immutable (cannot change), and physiological (mapped to biological sex).
Approach: They examine female first author percentages and citations to papers in natural language processing . they find that only about 29% of first authors are female and only about 25% of last authors are male .
Outcome: The authors show that only about 29% of first authors are female and only about 25% of last authors are male . the authors argue that gender gaps are unfair and need to be addressed .
MarathiEmoExplain: A Dataset for Sentiment, Emotion, and Explanation in Low-Resource Marathi (2025.findings-emnlp)

Copied to clipboard

Challenge: Marathi is the third most widely spoken language in India with over 83 million native speakers . available Marath datasets are limited to coarse sentiment labels and lack fine-grained emotional categorization or interpretability through explanations.
Approach: They propose to annotate Marathi sentences labeled with sentiment, emotion and a corresponding natural language justification.
Outcome: The proposed dataset provides a benchmark for future research in multilingual and explainable NLP.
Exploring the Role of Argument Structure in Online Debate Persuasion (2020.emnlp-main)

Copied to clipboard

Challenge: Existing work in NLP has shown that linguistic features extracted from debate text and features encoding the characteristics of the audience are both critical in persuasion studies.
Approach: They propose to incorporate argument structure features into an LSTM-based model to assess the persuasiveness of debates.
Outcome: The proposed model incorporates argument structure features to predict debaters that make the most convincing arguments on online debate forums.
NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for evaluating large language models using annotated benchmarks are in trouble . data contamination can cause wrong scientific conclusions being published .
Approach: They argue that the evaluation of NLP tasks using annotated benchmarks is in trouble . they define different levels of data contamination and propose a community effort .
Outcome: The proposed measures should detect when data from a benchmark was exposed to a model and flag papers with conclusions compromised by data contamination.
Indian Language Wordnets and their Linkages with Princeton WordNet (L18-1)

Copied to clipboard

Challenge: Wordnets are rich lexico-semantic resources. Linked wordnets link similar concepts in wordnet of different languages.
Approach: They propose to map 18 Indian wordnets linked with Princeton WordNet . they use expansion approach with Hindi Wordnet as pivot .
Outcome: The proposed mappings of 18 Indian wordnets are based on Princeton WordNet . they show that availability of such resources will have a direct impact on NLP progress .
Automatic Gloss-level Data Augmentation for Sign Language Translation (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods for enhancing sign language text data are insufficient . fewer studies have been performed on text data augmentation compared to video data .
Approach: They propose three methods to augment sign language text data using Korean sign language gloss dictionary.
Outcome: The proposed method improves translation performance by 0.204 and 0.170 compared to the original data.
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
Enhancing Idiomatic Representation in Multiple Languages via an Adaptive Contrastive Triplet Loss (2024.findings-acl)

Copied to clipboard

Challenge: Accurately modeling idiomatic or non-compositional language has been a longstanding challenge in natural language processing (NLP).
Approach: They propose an approach to model idiomaticity effectively using a triplet loss that incorporates the asymmetric contribution of components words to an idiomatic meaning by using adaptive contrastive learning and resampling miners.
Outcome: The proposed model outperforms previous models significantly on a SemEval challenge and outperformed previous alternatives in many metrics.
A Survey on Natural Language Processing for Fake News Detection (2020.lrec-1)

Copied to clipboard

Challenge: Automated fake news detection is a critical but challenging problem in NLP . social media has accelerated the spread of fake news, threatening public safety .
Approach: They describe the challenges involved in fake news detection and describe related tasks . they outline promising research directions and highlight the difference between fake news and related tasks.
Outcome: The proposed models are more fine-grained, detailed, fair, and practical.
Identifying Fine-grained Depression Signs in Social Media Posts (2024.lrec-main)

Copied to clipboard

Challenge: Currently, most studies focus on a binary classification setup or on pre-established resources.
Approach: They evaluated machine learning techniques to model 21 depression signs in social media posts from Brazilian undergraduate students.
Outcome: The proposed methods struggle to classify the majority of depression signs on social media posts, compared with the majority on the social media sites.
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Traditional approaches to truncate inputs, sparse self-attention, and chunking often lead to information loss and hinder the model’s ability to capture long-range dependencies.
Approach: They propose a novel chunk representation method that uses unsupervised keyphrase extraction to group input tokens to retain core document content while reducing input length.
Outcome: The proposed method minimizes information loss and improves the efficiency of Transformer-based models.
Can Transformers Reason in Fragments of Natural Language? (2022.emnlp-main)

Copied to clipboard

Challenge: Recent work on natural language inference has identified two strands of research .
Approach: They investigate whether neural networks have acquired logical principles from natural language . they use transformer-based models to detect valid inferences in controlled fragments of natural language.
Outcome: The proposed model overfits to superficial patterns in the data rather than acquiring the logical principles governing reasoning in natural language fragments.
TokenDrop + BucketSampler: Towards Efficient Padding-free Fine-tuning of Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Pre-training of Language Models (LMs) is a challenge due to its huge computational footprint.
Approach: They propose a framework that improves the efficiency and accuracy of LM fine-tuning by removing padding tokens from sequences that are variable-length .
Outcome: The proposed framework accelerates fine-tuning on diverse downstream tasks by 10.61X while producing models that are up to 1.17% more accurate compared to conventional fine-uning.
Automatic Period Segmentation of Oral French (2020.lrec-1)

Copied to clipboard

Challenge: Analor is a semi-automatic tool for speech segmentation in periods but it only takes into account prosodic characteristics of speech.
Approach: They propose to use a Fribourg model of macro-syntax to detect periods in syntactic and prosodic terms to develop an automatic tool for automatic segmentation of linguistic units.
Outcome: The proposed tool is compared with an existing tool Analor which divides speech into smaller segments and that CRF models detect larger segments rather than macro-syntactic periods.
IndicFinNLP: Financial Natural Language Processing for Indian Languages (2024.lrec-main)

Copied to clipboard

Challenge: IndicFinNLP is a collection of 9 datasets relating to FinNLP for three Indian languages.
Approach: They propose to use financial NLP to detect exaggerated numerals in financial texts written in Hindi, Bengali, and Telugu.
Outcome: The proposed framework detects exaggerated numerals in financial texts written in Hindi, Bengali, and Telugu.
“So You Think You’re Funny?”: Rating the Humour Quotient in Standup Comedy (2021.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for humour classification are limited due to the subjectivity of the content and the multiple interpretations of the data.
Approach: They propose to annotate a multi-modal humour-annotated dataset using stand-up comedy clips and compute a humor quotient using the audience's laughter.
Outcome: The proposed scoring mechanism is validated by comparing with manual scoring methods and achieves an accuracy of 0.813 in terms of QWK.
We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields (2023.emnlp-main)

Copied to clipboard

Challenge: In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other)
Approach: They quantify the degree of influence between 23 fields of study and NLP on each other . they find that cross-field engagement of NLP has declined from 0.58 in 1980 to 0.31 in 2022 .
Outcome: The proposed Citation Field Diversity Index (CFDI) has declined from 0.58 in 1980 to 0.31 in 2022, the authors show .
The Less the Merrier? Investigating Language Representation in Multilingual Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Multilingual models can be used to integrate multiple languages into one model and use cross-language transfer learning to improve performance for different NLP tasks.
Approach: They propose to include languages in popular multilingual models and to use cross-language transfer learning to improve performance for different NLP tasks.
Outcome: The proposed models perform better on downstream tasks for seen and unseen languages than community-centered models for low-resource languages.
A French Corpus for Semantic Similarity (2020.lrec-1)

Copied to clipboard

Challenge: Semantic textual similarity is a subtask of Natural Language Processing.
Approach: They propose to use an annotation corpus for French to assess semantic similarity . they use an annotated corpus with 1,010 sentence pairs with five annotators .
Outcome: The proposed corpus for French is the first that we know of.
DiffusionSL: Sequence Labeling via Tag Diffusion Process (2023.findings-emnlp)

Copied to clipboard

Challenge: Sequence Labeling (SL) is a long-standing field of natural language processing.
Approach: They propose a framework that utilizes a conditional discrete diffusion model for generating discrete tag data.
Outcome: The proposed framework outperforms gpt-3.5-turbo on multiple benchmark datasets and tasks.
Efficient Data Learning for Open Information Extraction with Pre-trained Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Experimental results indicate that, compared to previous SOTA methods, OK-IE requires only 1/100 of the training data (900 instances) and 1/120 of the time (3 minutes) to achieve comparable results.
Approach: They propose a framework that transforms OpenIE into the pre-training task form of the T5 model, thereby reducing the need for extensive training data.
Outcome: The proposed framework transforms OpenIE into the pre-training task form of the T5 model, reducing the need for extensive training data and significantly reducing training time.
Subspace Chronicles: How Linguistic Information Emerges, Shifts and Interacts during Language Model Training (2023.findings-emnlp)

Copied to clipboard

Challenge: Contemporary advances in NLP are built on the representational power of latent embedding spaces learned by self-supervised language models (LMs).
Approach: They use a new information theoretic probing suite to analyze representational subspaces in language models.
Outcome: The proposed approach compared performance of nine tasks across 2M pre-training steps and five seeds.
Graphically Speaking: Unmasking Abuse in Social Media with Conversation Insights (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to detect abusive language often ignore conversational context, leading to inconsistent and sometimes inconclusive results.
Approach: They propose a graph neural network approach that uses conversational context to model social media conversations as graphs, where nodes represent comments and edges capture reply structures.
Outcome: The proposed model outperforms baseline and linear context-aware methods and achieves significant improvements in F1 scores.
BioDEX: Large-Scale Biomedical Adverse Drug Event Extraction for Real-World Pharmacovigilance (2023.findings-emnlp)

Copied to clipboard

Challenge: pharmacovigilance (PV) is a tool for analyzing adverse drug events from biomedical literature . pharmacologists use natural language processing to extract core information from papers .
Approach: They propose a resource for biomedical adverse drug event eXtraction using natural language processing.
Outcome: The proposed model achieves 59.1% F1 (validation) and estimates human performance to be 72.0% F1 . the proposed model could be used to improve drug safety monitoring, also called pharmacovigilance, in the future.
Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on metaphor processing have focused on single datasets and specific task settings, often using artificially constructed data through lexical replacement.
Approach: They propose to evaluate the capabilities of Large Language Models (LLMs) in metaphor interpretation across multiple datasets, tasks, and prompt configurations.
Outcome: The proposed frameworks are more realistic and efficient than current models and are more efficient than existing models.
Leveraging Linguistically Enhanced Embeddings for Open Information Extraction (2024.lrec-main)

Copied to clipboard

Challenge: Open Information Extraction (OIE) is a structure prediction task in NLP that aims to extract structured n-ary tuples from free text.
Approach: They propose to leverage linguistic features with a Seq2Seq PLM for OIE to improve performance.
Outcome: The proposed methods give any neural OIE architecture the key performance boost from both PLMs and linguistic features in one go.
M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: a new ensemble decoding approach enhances the performance of Large Language Models.
Approach: They propose a multi-prompt ensemble decoding approach to enhance LLM performance . they submit n variations of prompts with X to LLMs in batch mode to decode and derive probability distributions .
Outcome: The proposed method improves pass@k rates, LENS metrics and BLEU scores on diverse NLP tasks.
HANSEN: Human and AI Spoken Text Benchmark for Authorship Analysis (2023.findings-emnlp)

Copied to clipboard

Challenge: Authorship Analysis is an essential aspect of Natural Language Processing (NLP) for a long time.
Approach: They propose to use 17 human speech datasets and 3 LLMs to create a benchmark for spoken texts.
Outcome: The proposed benchmark encompasses 17 human datasets and AI-generated spoken texts created using 3 prominent LLMs: ChatGPT, PaLM2, and Vicuna13B.
Annotating the Annotators: Analysis, Insights and Modelling from an Annotation Campaign on Persuasion Techniques Detection (2025.findings-acl)

Copied to clipboard

Challenge: Existing annotation campaigns based on heuristic guidelines have not been thoroughly discussed.
Approach: They propose a probabilistic model for optimizing intervention scheduling to reduce the cost of an expert oversight in annotation tasks.
Outcome: The proposed model advocates for an expert oversight in annotation tasks and periodic quality audits to reduce costs.
CASE: Efficient Curricular Data Pre-training for Building Assistive Psychology Expert Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to identify mental health disorders rely on limited availability of psychologists.
Approach: They propose to use forum posts to analyze text data to identify mental health issues . they propose to utilize readily available curricular texts for pre-training pipelines .
Outcome: The proposed pipelines achieve an f1 score of 0.91 for Depression and 0.88 for Anxiety compared to existing pipelines.
Evaluating Gender Bias of LLMs in Making Morality Judgements (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable capabilities in a multitude of NLP tasks, but are still not immune to limitations such as gender bias.
Approach: They propose to use a dataset to examine whether LLMs possess gender bias when asked to give moral opinions.
Outcome: The proposed models show that they are biased when asked to give moral opinions.
LoSST-AD: A Longitudinal Corpus for Tracking Alzheimer’s Disease Related Changes in Spontaneous Speech (2024.lrec-main)

Copied to clipboard

Challenge: Language-based biomarkers have shown promising results in differentiating those with Alzheimer’s disease (AD) diagnosis from healthy individuals, but the earliest changes in language are thought to start years or even decades before the diagnosis.
Approach: They propose to use transcripts of public interviews with 20 famous figures to track language change over several decades to validate their corpus.
Outcome: The proposed corpus can provide a valuable starting point for the development of early detection tools and enhance our understanding of how AD affects language over time.
Meta-Evaluation of Sentence Simplification Metrics (2024.lrec-main)

Copied to clipboard

Challenge: Automatic Text Simplification (ATS) is a major natural language processing task that aims to help people understand complex text.
Approach: They propose to use a human-annotated dataset to study automatic text simplification models to determine which metrics to use when evaluating new models.
Outcome: The proposed models reconstruct the text into a simpler format by deletion, substitution, addition or splitting, while preserving the original meaning and correct grammar.
The Zeno’s Paradox of ‘Low-Resource’ Languages (2024.emnlp-main)

Copied to clipboard

Challenge: 'low resource' languages are understudied by the NLP community, while 'high resource' is referred to as 'achieved', while high-resource languages are referred .
Approach: They qualitatively analyzed 150 papers from the ACL Anthology and popular speech-processing conferences that mention the keyword ‘low-resource.
Outcome: The proposed analysis reveals that several interacting axes contribute to ‘low-resourceness’ of a language and why that makes it difficult to track progress for each individual language.
Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to model complex subjective tasks in natural language are limited by significant variation in annotations.
Approach: They propose a simple in-context learning binary filtering baseline that estimates the reasonableness of a document-label pair.
Outcome: The proposed approach can be integrated into annotation pipelines to enhance signal-to-noise ratios.
Stochastic Fine-Tuning of Language Models Using Masked Gradients (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are the dominant paradigm in Natural Language Processing but fine-tuning them for specific downstream tasks often requires updating a vast number of parameters.
Approach: They propose a method that selectively updates a small subset of parameters in each step of the tuning process.
Outcome: The proposed approach outperforms existing fine-tuning methods while updating merely **0.08**% of the model’s parameters.
Multilingual Brain Surgeon: Large Language Models Can Be Compressed Leaving No Language behind (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for MC focus on quantization and network pruning.
Approach: They propose a calibration method that samples calibration data from various languages proportionally to the language distribution of the model training datasets.
Outcome: The proposed method improves the performance of existing English-centric compression methods on the BLOOM multilingual LLM.
TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling (2025.findings-emnlp)

Copied to clipboard

Challenge: TransBERT framework for pre-training language models using exclusively synthetically translated text is limited in specialized domains.
Approach: They propose a framework for pre-training language models using exclusively synthetically translated text . they also introduce a scalable translation toolkit that leverages synthetically trained data .
Outcome: The proposed framework can be used to train language models using synthetically translated text . transCorpus toolkit can be scalable to the life sciences domain in french .
New Evaluation Methodology for Qualitatively Comparing Classification Models (2024.lrec-main)

Copied to clipboard

Challenge: Text Classification is one of the most common tasks in Natural Language Processing.
Approach: They propose a method for performing qualitative assessment over multiple classification models using a fine-tuned BERT and Logistic Regression evaluation methodology.
Outcome: The proposed evaluation methodology outperforms the baseline model in linguistic clustering and Sentiment Analysis.
Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are vulnerable to backdoor attacks, where adversaries poison a small subset of data to implant hidden behaviors.
Approach: They propose a training pipeline that immunizes instruction-tuned LLMs against backdoor attacks.
Outcome: The proposed defenses lower attack success rates while preserving instruction-following ability.
QUEEREOTYPES: A Multi-Source Italian Corpus of Stereotypes towards LGBTQIA+ Community Members (2024.lrec-main)

Copied to clipboard

Challenge: a dataset of social media texts addressing LGBTQIA+ individuals is presented in this paper . the dataset is based on two sources in italian: Facebook and Twitter .
Approach: They describe a dataset composed of two sub-corpora from two different sources in Italian . the dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events .
Outcome: The QUEEREOTYPES dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events.
RAAMove: A Corpus for Analyzing Moves in Research Article Abstracts (2024.lrec-main)

Copied to clipboard

Challenge: RAAMove is a comprehensive multi-domain corpus dedicated to the annotation of move structures in Research Article (RA) abstracts.
Approach: They propose a multi-domain corpus dedicated to the annotation of move structures in RA abstracts.
Outcome: The proposed corpus is based on a human-annotated dataset and a BERT-based model to verify its effectiveness.
In Benchmarks We Trust ... Or Not? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for Large Language Models (LLMs) are inadequate and lack a clear solution.
Approach: They propose checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage.
Outcome: The proposed checklists cover all aspects of benchmarking issues, both for benchmark creation and usage.
The Nature of NLP: Analyzing Contributions in NLP Papers (2025.acl-long)

Copied to clipboard

Challenge: despite this, what constitutes NLP research remains debated .
Approach: They propose a taxonomy of research contributions and introduce a task of automatically identifying contribution statements and classifying their types from NLP research papers.
Outcome: The proposed model analyzes 29k NLP research papers to understand their contributions .
LLMs as Planning Formalizers: A Survey for Leveraging Large Language Models to Construct Automated Planning Models (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models excel in various natural language tasks but struggle with long-horizon planning problems requiring structured reasoning.
Approach: They propose to integrate large language models into AP and NLP planning frameworks by reviewing current research and identifying critical challenges and future directions.
Outcome: The proposed frameworks are used to support reliable off-the-shelf AP planners.
SlovakSum: A Large Scale Slovak Summarization Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets with hundreds and thousands of documents are mainly in the English language, but the available data is small or non-existent.
Approach: They propose to use a large Slovak news summarization dataset to evaluate its performance . the dataset contains headlines, short abstracts, and full source text .
Outcome: The proposed dataset is compared with a standard ROUGE metric and a mT5 model to evaluate its performance.
GMSA: Enhancing Context Compression via Group Merging and Layer Semantic Alignment (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved remarkable performance across NLP tasks . however, in long-context scenarios, they face high computational cost and information redundancy.
Approach: They propose an encoder-decoder context compression framework that generates a compact sequence of soft tokens for downstream tasks.
Outcome: Experiments show that GMSA outperforms baselines on multiple long-context question answering and summarization benchmarks while maintaining low end-to-end latency.
Strengthening the WiC: New Polysemy Dataset in Hindi and Lack of Cross Lingual Transfer (2024.lrec-main)

Copied to clipboard

Challenge: a new study addresses the problem of natural language processing in low-resource languages such as Hindi . the paper focuses on Word Sense Disambiguation, a fundamental NLP task that deals with polysemous words.
Approach: They propose a Hindi WSD dataset that allows training and testing of contextualized models.
Outcome: The proposed dataset enables training and testing of contextualized models in Hindi . the results show that the proposed dataset can handle polysemy tasks in low-resource languages .
Target-Adaptive Consistency Enhanced Prompt-Tuning for Multi-Domain Stance Detection (2024.lrec-main)

Copied to clipboard

Challenge: Stance detection is a fundamental task in natural language processing, but it is challenging due to diverse expressions and topics related to the targets from multiple domains.
Approach: They propose a prompt-tuning method that incorporates target knowledge and prior knowledge to construct target-adaptive verbalizers for diverse domains.
Outcome: The proposed method outperforms the state-of-the-art methods on nine stance detection datasets from multiple domains.
The Challenges of Creating a Parallel Multilingual Hate Speech Corpus: An Exploration (2024.lrec-main)

Copied to clipboard

Challenge: Hate speech is one of the most demanding topics in Natural Language Processing, as its multifaceted nature is accompanied by a handful of challenges, such as multilinguality and cross-linguality.
Approach: They propose a pipeline that could be used to create a parallel multilingual hate speech dataset using machine translation.
Outcome: The proposed pipeline will be able to create a parallel multilingual hate speech dataset using machine translation.
WkNER: Enhancing Named Entity Recognition with Word Segmentation Constraints and kNN Retrieval (2024.lrec-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) tasks require detecting the span and category of the entity from the text block.
Approach: They propose a kNN retrieval enhancement algorithm that incorporates word segmentation information to enhance the model’s generalization ability and alleviate the problem of missing entity tokens in prediction.
Outcome: The proposed method improves the performance of baseline models and achieves better or compared recognition accuracy than previous state-of-the-art models in multiple public Chinese and English datasets.
A Systematic Survey of Automatic Prompt Optimization Techniques (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in prompt engineering have created impediments for end users to adopt . however, prompt engineering remains an impedance due to rapid advances in models, tasks, and associated best practices.
Approach: They propose to define APO as a 5-part unifying framework and categorize all relevant works based on their salient features.
Outcome: The proposed framework aims to improve the performance of large language models on various tasks.
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos (2025.emnlp-main)

Copied to clipboard

Challenge: This paper explores using Multimodal Large Language Models (MLLMs) to respond to student questions from online lectures . MLLM is a novel question answering task of real world significance .
Approach: They propose to use Multimodal Large Language Models to automatically respond to student questions from online lectures by using a dataset of 5252 question-answer pairs from 296 computer science videos.
Outcome: The proposed model can fine tune and fine tune questions from 296 computer science videos and show that students' preferences are important to the task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations